Compare commits

...

16 Commits

Author SHA1 Message Date
admin 18d03bd437 v0.132.0: the slow crash-loop counter (R-539, operator ruling 3 of 2026-09-16)
gates / gates (push) Successful in 12s
Beside the unchanged 3-in-15-minutes brake, a second counter: restarts the
supervisor performed in the last 24 hours. At the fifth the heartbeat stanza
sets slow_crashloop_since (moving at most once per 24 h), slow_crashloop and
restarts_24h; hub v0.117.0 mints controller_slow_crashloop (warning,
operator-only) when the timestamp moves. It never stops restarting.

Persisted per guest (tmp+rename, 0600) so an agent restart or reboot does not
reset it - unlike the fast record, whose reason for staying in memory (a
persisted give-up outliving the fix) does not apply to a counter that only
warns. Deliberate kills count. The startup line prints the new limits.

Red-proofs seen failing: no counter; the once-per-24h guard removed ('the
operator would be mailed per restart'); the save removed ('Restarts24h:1'
after an agent restart). Negative control: restarts 7 h apart never raise it.
go build/vet/test ./... green, 30 packages.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 10:26:42 +02:00
admin e98b857684 REPORT: v0.131.0 supervisor + per-tier status, delivery and live validation
gates / gates (push) Successful in 13s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 11:37:07 +02:00
admin dcdeb3d16d CHANGELOG: v0.131.0 released (tag + package verified by download)
gates / gates (push) Successful in 12s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 09:50:47 +02:00
admin 610804b98d v0.131.0: controller supervisor (R-523); per-tier backup status + tier storage presence (R-517/R-518)
gates / gates (push) Successful in 11s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 09:49:44 +02:00
admin 4586f0f7f6 re-run CI against a register that now carries R-421
gates / gates (push) Successful in 9s
The earlier run convicted correctly: instructions_gate found this repo citing R-421 while
felhom.eu's OPEN-ITEMS.md did not yet have the row. My ordering, not the gate's fault - the register
lives in felhom.eu, so a repo citing a new row must be pushed after it.
2026-09-01 12:45:52 +02:00
admin 205e22babe decoy sweep: no gate changed here, and that is the result (R-421)
gates / gates (push) Failing after 12s
All 29 gate scripts across the four repos were read and DECOYED - the label constructed without the
fact, the gate run, the verdict recorded. 16 were fooled. None of them were in this repo.

A decoy that nobody would write proves nothing, so the attempts that turned out illegitimate were
WITHDRAWN rather than counted. Both of this repo were withdrawn, and both are named in the audit.

The gates here that could not be given a plausible decoy are listed BY NAME in
felhom.eu/scripts/decoy_coverage_gate.py EXEMPT (R-426) as UNTESTED - not as sound. A gate nobody
tried to fool is UNKNOWN, and calling it sound would be the same confident guess this sweep exists
to find.

Survey table: felhom.eu/documentation/audits/AUDIT-gate-decoys-2026-09-01.md
2026-09-01 12:38:58 +02:00
admin 058b945064 gate 11: register the shared observations gate
gates / gates (push) Successful in 10s
felhom.eu/scripts/observations_gate.py, invoked across the workspace like
reuse_refs_check.py and instructions_gate.py. This repo's REPORT.md has no
observations section today, so the gate passes quietly - it is registered for
the session that writes one.
2026-08-23 13:53:20 +02:00
admin 40d857b527 CHANGELOG: the fleet runs the published bytes, not the proof build (R-349)
gates / gates (push) Successful in 7s
Both boxes were first given a hand build: same source, same version
string, different bytes (256e0829 vs the published a56a92a7), because
release-agent.sh builds with -trimpath -buildvcs=false and a hand build
does not.

Nothing would have corrected it. The boxes already reported 0.130.0, so
self-update saw the vouched version as installed and would have done
nothing, forever. Every version check in the system compares the STRING.

Both reinstalled from the downloaded package; both now report a56a92a7.
2026-08-20 12:53:37 +02:00
admin 7ae6990bac v0.130.0 released: tag + package published, heading now claims it
gates / gates (push) Successful in 8s
sha256 a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3
tag v0.130.0 at 7569f34, 14,141,158 bytes.

Reproducible: rebuilding with -trimpath -buildvcs=false matches the
published artifact byte for byte (R-186's property, checked not assumed).

The heading said UNRELEASED while the fix was hand-installed on demo-hp
only -- publishing then would have pushed it onto the control box through
self-update. release-complete convicted on the release heading and was
right to; the answer was to stop claiming a release, not to bypass it.

NOT VOUCHED by this commit. Vouching is the separate operator act.
2026-08-20 12:47:25 +02:00
admin 7569f34aeb CHANGELOG: correct a claim this session made and then disproved
gates / gates (push) Successful in 8s
The v0.130.0 draft said the 388 descriptors already stuck on ep0 would
persist until the PBS proxy restarted. Measured within the hour: they
clear when the AGENT restarts. demo-hp released its 199 in one second
(415 -> 216 fd); demo-felhom released the remaining 203 (220 -> 17 fd in
under two seconds). 17 is precisely ep0's t0 baseline of 2026-08-18.

They were held on both sides. Closing either side ends them. ep0 was
read-only throughout and its proxy PID never changed.
2026-08-20 12:39:12 +02:00
admin ede49b610d R-344: restore the idle-connection timeout our hand-rolled transports lost
gates / gates (push) Successful in 7s
Every client here pins TLS, so none can use http.DefaultTransport and each
hand-rolls its own. A composite literal takes IdleConnTimeout ZERO, which
means retain idle connections forever -- not "use a sane default".
pbsTargetsFromPVE builds a fresh pbs.Client every cycle and drops the
previous one, and an abandoned http.Transport does not close its
connections. One stranded socket per cycle, on both sides, forever.

Measured: 388 established connections on ep0 over 46 h, 194 per box, zero
closed in a 31-minute window. pvestatd and proxmox-backup-client made
162,404 requests in the same window and leaked none.

New leaf package internal/httpx owns the default (90s, http.DefaultTransport's
own value) and NewTransport, which returns a FRESH transport per call and
treats <=0 as "use the default", never "no timeout". pbs.Config gains
IdleConnTimeout for tests only.

hub and proxmox carried the same missing default and are corrected here for
consistency. Neither contributed to the ep0 leak -- both are built once per
process and neither talks to ep0:8007.

Tests count connections SERVER-side and model the abandonment, so they pin
the consequence rather than the field. Two red-proofs, both seen failing:
removing the timeout -> "still holds 5 open connection(s), want 0";
DisableKeepAlives -> "3 sequential requests over 3 connection(s), want 1"
(the leak test PASSES under that one -- it is the worse-fix guard that
catches it).

Not released: hand-installed on demo-hp only so demo-felhom stays the
control. CHANGELOG heading stays UNRELEASED until the publish is authorised.
2026-08-20 11:09:39 +02:00
admin f17ed11599 REPORT: agent v0.129.0 — the retained-package recovery class, released and deployed
gates / gates (push) Successful in 21s
2026-08-12 18:49:51 +02:00
admin 1db56bf837 v0.129.0 — a correct code for an earlier package stops being called wrong (R-311)
gates / gates (push) Successful in 14s
Yesterday's drill proved a retained escrow package opens a set-aside store and
restores planted files byte-identical, while this agent answered the customer's
correct code with "the recovery code did not open the sealed bundle". Nothing had
ever tried the retained packages, so a correct-but-earlier code and a mistype were
genuinely indistinguishable.

OffsiteKeyRecoverer gains an optional FetchRetained, consulted ONLY after the
current package refuses, so the ordinary recovery pays nothing for it and cannot
fail because of it. A match returns ErrCodeOpensRetained wrapped in a
RetainedOpenedError carrying the supersession date - no material, no code, no
password. The local API answers 422: a FIFTH status added to the R-224 switch,
never a restructuring of it.

Fail-safe in every direction. Nil fetcher, a hub too old for the route (404 is a
clean "none"), a transport failure, a malformed package: each leaves the original
refusal standing. Attempts bounded at 6 because each unwrap is ~1s of scrypt.

Seven tests with REAL age crypto - the two situations are indistinguishable AT
THE UNWRAP, so a faked unwrap would prove nothing. Red-proof asserted applied:
remove the retained lookup and the fail-closed wrong-code error returns, which is
the lie in those exact words.
2026-08-12 18:40:00 +02:00
admin 53d047a6c1 Two guards, one number: bound the published check to the retention it must live with
gates / gates (push) Successful in 17s
Gates only. No release, no version bump, no binary published; the agent stays
v0.128.0 at 28ba8593b8 and nothing on a customer's machine changes.

THE COUPLING DEFECT. The registry stopped serving 0.120.0 and older while
check-published-versions.py demanded every tag still be downloadable. Both rules
are sensible and together they are impossible, so CI went red at a commit whose
own run had been GREEN the day before -- and would have gone red again at the
next publish when 0.121.0 was evicted. scripts/retention-policy.json is now THE
number and both readers take it from there.

WHAT CI NO LONGER COVERS, and it prints this on every run rather than leaving it
to be discovered: a released version older than the retention window is no longer
asserted downloadable. Its git tag and its config tree ARE still asserted -- only
the binary's presence is dropped. A missing policy file is INCONCLUSIVE (exit 2),
never silently unbounded.

THE NUMBER IS NOT A LOCATED RULING and the file says so in its own header. Ten is
what the registry demonstrably holds; no register row records a prune, R-210 says
"Nothing was deleted; this is a list, not an action" and concerns local Docker
images, and container packages hold 19 each. The principled bound is the hub's
vouched min_agent floor -- nothing can install below it -- and that is the
recorded follow-up.

check-release-complete.py is the tag half as a machine. release-agent.sh already
warned that "a released version without a git tag 404s a box mid-install, as
root" and the step was still missed, so this is a gate and not a reminder. Legs
1-2 need no network and run in --fast, so the pre-push hook is the earliest
catch. Red-proved by repointing the CHANGELOG head at an unreleased v0.129.0:
both legs convicted and each named its fix command.

Three controls run: green at 10 naming what it dropped; widened to 11 the evicted
version re-enters and convicts; policy removed gives INCONCLUSIVE naming the path.
2026-08-09 19:05:05 +02:00
admin 28ba8593b8 v0.128.0 — R-221: the escrow seed is asserted every tick, not remembered once
gates / gates (push) Failing after 28s
A rebuilt box could not run the escrow ceremony AT ALL, with no way forward from inside the product.
This was the only open item blocking a customer from something we promise them.

MECHANISM, established at file:line rather than assumed. The preflight refuses on
escrow.pbs_storage_id; the pbsdr bridge writes that key; and it wrote it from exactly one place —
finishConverged, reached only on the paths that actually converge.

The marker and the key live in different places and die at different times. The marker is host-side
(<agent-state>/pbsdr/marker.json). The key is in agent.json, which step_agent_config renders from
`base = {}` unless an explicit --preserve-from is given (felhom.eu/scripts/felhom-host-install.sh:
2396 the step, :2449 the render, :2579 the O_TRUNC write; the flag :1246, defaulting empty at :256)
— AND THE RENDER NEVER WRITES AN escrow SECTION AT ALL (grep over the whole heredoc: zero hits). So
a rebuild keeps the marker and takes the key: same descriptor, same hash, early return, and the seed
never runs again into a config that no longer has it.

A rebuild is only the case that was measured. The same hole opens for a hand-edited or restored
config, which is the honest reason this is a seam fix rather than an installer fix: the seed must be
a thing the loop ASSERTS, not a thing it did once.

Apply now re-asserts the seed BEFORE the idempotent early return. seedEscrowStorageID is unchanged
and still never clobbers a different existing value — an operator's own choice outranks the
descriptor's, with a warning naming both.

THE EARLY RETURN IS KEPT. It stops a converged box re-running Proxmox operations every 60s, and
TestSeedReasserted_OnConvergedTick_WithZeroProxmoxCalls asserts ZERO recorded runner calls on that
tick, so a "fix" that simply deleted the return would fail. Cost: one small file read plus a JSON
parse per tick, no exec, no network, early-returning once the value matches.

A seed failure can never un-converge the box: Warn plus a message on the published status, exactly
as finishConverged does it — no marker write, no state change.

Tests drive the REAL Apply with a real temp-dir agent.json and a call-recording runner; calling
seedEscrowStorageID directly cannot see the early return, which IS the defect. Production wiring
(pbsdr.NewManager(..., cfg.SourcePath, ...)) is asserted by walking main.go's AST, not by
strings.Contains, which a commented-out call also satisfies.

Red-proofs, each with the mutation asserted applied: removing the new call makes Scenario A fail on
today's tree (it did, with the intended message); removing the early return makes the
zero-Proxmox-calls assertion fail (it did).

go build / go vet / go test ./... green (29 packages), run separately from this commit.
2026-08-08 16:29:13 +02:00
admin 6981450110 docs: a comment claimed the hub reads a field it has no field for (R-260)
gates / gates (push) Successful in 26s
Comment-only; no behaviour, no wire change, no version bump, nothing to rebuild.

HostReport.SelfUpdatePending / SelfUpdatePendingVersion carried "The hub reads an absent field as
pending=false, the correct default." The hub has NO FIELD for either, so it reads nothing — present
or absent — and encoding/json discards them on arrival. The sentence described an intent rather than
the code and read as settled long enough that a class sweep had to find it.

The emission is correct and stays: the agent reports the truth and the fault is entirely in the
receiving. The missing consumer is R-264 (OPEN). felhom.eu/scripts/wire_contract_gate.py now refuses
any NEW field of this shape and records the existing ones as reasoned allowlist entries.
2026-08-08 08:46:17 +02:00
27 changed files with 2727 additions and 62 deletions
+285
View File
@@ -1,3 +1,286 @@
## Unreleased (to become v0.132.0) — a controller that dies slowly is reported, not just restarted (2026-09-17, R-539)
**MinAgent impact:** none required by any controller. Hub **v0.117.0** turns the new fields into
`controller_slow_crashloop`; an older hub ignores them.
- **R-539 (operator ruling 3 of 2026-09-16) — the slow crash-loop counter.** The 3-restarts-in-15-minutes
brake cannot see a controller that dies every 20 minutes (measured 2026-09-16, R-531: four restarts,
none accumulating, only an `info` event that mails nobody). Beside it, unchanged, a second counter:
restarts the supervisor performed in the last **24 hours**; at the **fifth**, the heartbeat's
`controller_supervisor` stanza sets `slow_crashloop_since` (and `slow_crashloop: true`,
`restarts_24h`). The hub mails on that timestamp MOVING; it moves **at most once per 24 hours**. It
does **not** stop restarting — the fast brake remains the only brake.
- **Persisted per guest** at `/var/lib/felhom-agent/guests/<vmid>/controller-restarts-24h.json`
(tmp + rename, 0600), so an agent restart or a host reboot does not reset it. The fast record stays
in memory; the reason it does (a persisted "give up" could outlive the fix) does not apply to a
counter that only warns. Unreadable or corrupt → WARN and a clean start, never a blocked supervisor.
- **Deliberate kills count.** The supervisor cannot tell an operator's `docker kill` from a crash
(measured 2026-09-15); a controller killed five times a day is worth a line either way.
- The supervisor's startup line now prints `slow_crashloop_max=5 slow_crashloop_window=24h0m0s`.
**Red-proofs, each seen failing:** five restarts 20 minutes apart raise it
(`TestControllerSupervisor_SlowCrashloop` — fails with no counter; fails again with the once-per-24-hours
guard removed, "the operator would be mailed per restart"); the counter survives an agent restart
(`…SlowCounterSurvivesAgentRestart` — fails with the save removed, `Restarts24h:1`); restarts seven hours
apart never raise it (the negative control). Wire shape extended with `restarts_24h` and `slow_crashloop`.
## v0.131.0 — a dead controller comes back by itself; the backup status speaks per tier (2026-09-15, R-523 / R-517 / R-518)
> **RELEASED 2026-09-15** by `scripts/release-agent.sh` — tag `v0.131.0`, sha256 `1118b552f7e775fbde9544c7764ede7e6046e0a7db16ae8494d07a18e3c2ac9c`. Delivered to boxes by the controller v0.243.0 floor (declared MinAgent), not by hand.
- **R-523 (P1) — the in-guest controller supervisor.** BIGNIGHT F9: `docker kill felhom-controller`
left the household's dashboard on 502 for 33 minutes, because nothing watched the container.
Measured first (2026-09-15, Docker 29.8.0, `evidence-p1fixes-2026-09-15/A1`): after `docker kill`,
BOTH `--restart unless-stopped` and `--restart always` leave the container `exited (137)` after
60 s — a policy change alone is not a fix. New `internal/localapi/controllersupervisor.go`: every
30 s, for each felhom-pool guest the agent provisioned (`<guests>/<vmid>/bootstrap` exists) that is
running, it reads `docker inspect -f {{.State.Status}} felhom-controller`; on the SECOND consecutive
not-running (or absent) observation it runs `systemctl restart felhom-controller-bootstrap.service`
inside the guest — the swap's own restart, over the same GuestExecutor and the same two sudoers
grants. No new privilege. Guards, each pinned by a test: not during a controller swap (the swap's
in-flight flag); not when parked (`touch /var/lib/felhom-agent/guests/<vmid>/controller-parked` on
the HOST); not on a stopped, locked or vzdump-busy guest; not on an unknown docker answer; not on a
guest the agent did not provision; and **no thrash** — 3 restarts in 15 minutes stop the restarts
for 30 minutes. The record rides the host report as `controller_supervisor` (additive,
`omitempty`); hub v0.114.0 mints `controller_restarted_by_agent` (info) and `controller_crashloop`
(error), both operator-only. Red-proofs: without the restart call, the kill test fails at
"restarts=0"; without the backoff block, the crash-loop test fails at "restarted 10 times".
- **Golden script** (`configs/build-golden.sh`): the controller runs `--restart always` (covers a
Docker daemon restart after a manual stop — nothing more). **No golden baked here** (R-468); existing
boxes keep `unless-stopped` until their next golden and are covered by the supervisor.
- **R-517 (P1) — `GET /backup/status` speaks per tier.** The untargeted response gains `tiers[]`:
per tier the newest SUCCESSFUL backup (`last_success`, from the record, or from the tier's storage
after an agent restart — `last_success_source: storage`), the last attempt kept apart
(`last_attempt {started_at, success, error}`), and whether the tier's storage exists (`storage:
present|absent|unknown`). `GET /backup/tiers` gains the same `storage` field (R-518's cheap half: the
controller skips an absent tier). `unknown` is never `absent` — a storage view that cannot be read
must not skip a backup. Additive; the untargeted `.backup` keeps its meaning (pinned). Red-proof:
filling `last_success` from the newest ATTEMPT fails at "pbs tier reports a failed attempt as its
last success".
## the decoy sweep — can this gate be fooled by a label? (2026-09-01, R-421) — NOT A RELEASE
**No product code, no version bump, no image, no golden.** A scripts change is not a release.
Four times in one week a gate turned out to match a NAME instead of the thing it named — R-410 (a
`mkdir` turned the release gate green), R-400 (seven debug controls answering nothing), R-378 (a
status word inside a sentence), R-419 (a phrase inside prose, including prose saying the marker was
absent). **All four found by accident.** The gates enforce everything else here and were the one part
nothing had checked.
**All 29 gate scripts read and decoyed. 16 were fooled.** 10 fixed here, 4 left with rows
(R-422..R-425), 6 could not be given a plausible decoy and are named (R-426 group d).
**The largest single cause was mundane:** eight gates set their SCOPE with `os.listdir` (one level).
Green and correct today; blind the moment anyone adds `templates/partials/`. `mojibake` and
`docker-v` already used `os.walk`, caught the identical planted file, and are the control that
proves the cause was the listing rather than the decoy.
Full survey table, and the five decoys withdrawn as illegitimate (mine, named):
`documentation/audits/AUDIT-gate-decoys-2026-09-01.md`.
**In this repo:** no gate changed, and that is the result. `release-complete` was decoyed and is
SOUND — the sweep's attempt (a non-version heading on top of `CHANGELOG.md`) was WITHDRAWN as
illegitimate, because `HEAD_RE.search` scans the whole file and still names `v0.130.0`. The three
shared gates are covered from `felhom.eu`; the remaining two are named in the decoy-coverage
exemption list (R-426) as UNTESTED, not as sound.
## v0.130.0 — the agent was the one leaking connections onto the off-site box (2026-08-20, R-344)
> **RELEASED 2026-08-20**, on the operator's word, after the fix was proved on both boxes.
> `sha256 a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3`, 14,141,158 bytes,
> tag `v0.130.0` at `7569f34`. Reproducible: a rebuild with `-trimpath -buildvcs=false` matches the
> published artifact byte for byte (R-186's property, checked rather than assumed).
>
> The heading read `## UNRELEASED — v0.130.0 candidate` until this point, deliberately: while the fix
> was hand-installed on `demo-hp` only, publishing would have pushed it onto `demo-felhom` through
> self-update and destroyed the control the proof rested on. **`release-complete` convicted on the
> release heading and was right to** — the answer was to stop claiming a release, not to bypass the
> gate. See R-347.
>
> **The fleet runs these exact bytes.** Both demo boxes were first given a hand build made during the
> proof — same source, same version string, **different bytes** (`256e0829…`), because
> `release-agent.sh` builds with `-trimpath -buildvcs=false` and a hand build does not. Nothing would
> have corrected that: the boxes already reported `0.130.0`, so self-update saw the vouched version as
> installed and would have done nothing, forever. Both were reinstalled from the **downloaded package**
> and now report `a56a92a7…`. Filed as **R-349**, because every prove-then-publish train hits it.
**What was measured, before anything was changed.** Between 2026-08-18 09:51:22Z and 2026-08-20
08:02:13Z, ep0's PBS proxy accumulated **388 established connections** — 194 from each demo box, on a
proxy whose descriptor ceiling is 65536 and whose runway at that rate was ~323 days. The connections
were held open on **both** sides: ep0 showed 388 while the two boxes showed 194 + 194, at two separate
instants, with the same source ports on each side, and **not one closed in a 31-minute window**.
`ss -tnp` on the boxes named the holder: **`felhom-agent`**, 194 of 194 on each, one PID.
**`pvestatd` and `proxmox-backup-client` made 162,404 requests in that window and leaked zero.** They
are 99.5% of the traffic to that endpoint and 0% of the leak. The agent made 811 requests — of which
387 were `GET .../snapshots` — and leaked 388 sockets. One per call, within one.
**The defect, and it is two things compounding.**
- `internal/pbs/client.go` built its transport as a composite literal:
`&http.Transport{TLSClientConfig: tlsCfg}`. That takes **`IdleConnTimeout` zero, which does not mean
"use a sane default" — it means retain idle keep-alive connections FOREVER.**
`http.DefaultTransport` sets 90s; hand-rolling the transport (which every client here must do,
because they all pin TLS) silently discards it.
- `pbsTargetsFromPVE` (`cmd/felhom-agent/main.go`) builds **a fresh `pbs.Client` every cycle**, as its
own doc comment says, and drops the previous one. An abandoned `http.Transport` does **not** close
its connections — it becomes unreachable while its `persistConn` read-loop goroutine keeps the
socket alive. So each cycle stranded exactly one connection that nothing could ever close.
The cadences reconcile without fitting: a 900s live-snapshot collect (184.7 cycles in the window) plus
a 6-hour verify loop (7.7) predicts 192.4 per box against **194 observed**.
**The fix is one field, restored to the standard library's own value.** New leaf package
`internal/httpx` owns `DefaultIdleConnTimeout = 90 * time.Second` — 90s because that is what
`http.DefaultTransport` uses, so there is nothing invented here to justify or tune — and
`NewTransport(tlsCfg, idleConnTimeout)`, which returns a **fresh** transport (never shared: each caller
pins a different endpoint) and treats a zero or negative timeout as **use the default, never "no
timeout"**. `pbs.Config` gains an `IdleConnTimeout` field that production leaves unset; only tests set
it, to avoid a 90-second wait.
**`internal/hub/client.go` and `internal/proxmox/client.go` carried the identical missing default and
were corrected in the same pass — but neither contributed to the ep0 leak, and this entry must not be
read as three leaks having been found.** Both are built **once per process**, so they held one idle
connection for the life of the daemon rather than accumulating, and neither talks to ep0:8007.
**Tests, and what they deliberately do not assert.** `internal/pbs/client_leak_test.go` counts
connections **server-side** and models what `pbsTargetsFromPVE` actually does — build a client, use it
once, drop it on the floor — then asserts the connections go away. It does not assert `err == nil` and
it does not assert that some field holds some value; both were true of the leaking code.
- **Red-proof 1 (the fix):** removing `IdleConnTimeout` from `NewTransport` fails the test with
*"after 5s the server still holds 5 open connection(s), want 0 (5 dialled in total)"* — the count is
in the message, so the failure cannot be mistaken for a timeout with another cause. Reverted.
- **Red-proof 2 (the fix that would be worse than the bug):** setting `DisableKeepAlives: true` also
makes the leak vanish — by dialling fresh for every request, which on a box polling ~40,000 times a
day is strictly worse than what we started with. **The leak test PASSES under that mutation**;
`TestPBSClient_KeepAliveStillReuses` is what catches it, failing with *"3 sequential requests over 3
connection(s), want 1"*. Reverted.
**What this release does NOT do.** It does not reduce the poll rate (**R-336 stays open, but re-scoped
— it was never the cause of this leak**), it does not refactor `pbsTargetsFromPVE` to cache or reuse
clients (a one-line default restores the standard behaviour; a lifecycle refactor adds
cache-invalidation questions for no measurable gain), and it adds no `CloseIdleConnections` call.
**One sentence in this entry was written before the deploy and was WRONG, and it is corrected here
rather than quietly edited.** It read: *"does not clear the 388 descriptors already stuck on ep0 —
those persist until that proxy restarts."* **Measured: they clear the moment the AGENT restarts.**
Replacing the binary on `demo-hp` released exactly its 199 descriptors within one second
(415 → 216 fd), and replacing it on `demo-felhom` released the remaining 203 (**220 → 17 fd in under
two seconds**). **17 is precisely ep0's `t0` baseline** of 2026-08-18 09:51:22Z. ep0 was read-only
throughout and its proxy PID never changed. The accumulated leak was never ep0's to hold on to — it
was held on both sides, and closing either side ends it.
## v0.129.0 — a correct code for an earlier package stops being called wrong (2026-08-12, R-311)
**The measurement this fixes.** On 2026-08-12 a recovery code that provably opens a RETAINED package
— unsealed by hand, and it restored planted files byte-identical from a store the box itself could no
longer open — was answered by this agent with *"the recovery code did not open the sealed bundle"*.
The code was correct. Nothing had ever tried the retained packages, so the engine could not tell a
correct-but-earlier code from a mistype, and the screen said so out loud: a true sentence about our
own incuriosity, read by the customer as a statement about their code.
**`OffsiteKeyRecoverer` gains an optional `FetchRetained`.** It is consulted ONLY after the current
package has refused, so the ordinary recovery pays nothing for it and cannot fail because of it. When
one of the retained packages opens, the recoverer returns `ErrCodeOpensRetained` wrapped in a
`RetainedOpenedError` carrying the supersession date — no material, no code, no password.
**The local API answers 422** ("your code is correct, it belongs to an EARLIER sealed package") — a
FIFTH status added to the R-224 switch, not a restructuring of it. 422 rather than 400 because the
request was well-formed AND the credential valid; a 400 would put it in the same bucket as a mistype,
which is the defect.
**Fail-safe in every direction.** A nil fetcher, a hub too old to have the route (404 is a clean
"none"), a transport failure, a malformed package: each leaves the original refusal standing,
unchanged. The worst outcome of this feature breaking is the behaviour we had before it existed.
Attempts are bounded (`MaxRetainedTried`, default 6) because each unwrap is ~1 s of scrypt by design
and an unbounded loop would turn one wrong code into a minutes-long hang.
**New hub client call:** `FetchRetainedIdentityEscrow` → `GET /api/v1/hosts/<id>/escrow/retained`
(hub >= v0.103.0), self-scoped by the same per-host key.
Seven tests with REAL age crypto, because the two situations are indistinguishable AT THE UNWRAP and a
faked unwrap would prove nothing about what was broken. Red-proof, asserted applied: removing the
retained lookup returns the fail-closed wrong-code error — **the lie comes back, in those words.**
---
### Gates only — 2026-08-09 (no release, no version bump, no binary published)
**Two guards, both owed since the 2026-08-09 install outage (R-273/R-287). Nothing that runs on a
customer's box changed; `scripts/` only, and the agent stays v0.128.0.**
- **`scripts/retention-policy.json` — THE retention number, in one file.** The registry stopped
serving `felhom-agent` 0.120.0 and older while `check-published-versions.py` demanded that every
tag still be downloadable. Both rules are sensible; together they are impossible, and CI went red
at a commit whose own run had been green the day before. The check now **reads the number from
this file** and bounds its assertion to the newest N generic versions.
**What CI no longer covers, said plainly rather than left to be discovered:** a released version
older than the retention window is **no longer asserted downloadable**. Its git tag and its configs
are still asserted — only the binary's presence is dropped. The check **prints exactly which
versions it stopped covering** on every run, so the narrowing cannot become permanent by accident.
**The number is an OBSERVED state, not a located ruling** — see the file's own header and R-287.
A missing or unreadable policy file is **INCONCLUSIVE (exit 2), never silently unbounded.**
- **`scripts/check-release-complete.py` — the tag half, as a machine.** Asserts that the version at
the head of `CHANGELOG.md` is tagged, that the tag points into this history, and that its package
is published. `release-agent.sh` already warned about this in as many words and the step was still
missed on 2026-08-08, which is why this is a gate and not a reminder. Legs 1–2 need no network and
therefore run in `--fast`, so the pre-push hook catches a missing tag at the earliest moment.
Registered in `agent_gates.py`; red-proved by pointing the CHANGELOG head at an unreleased
v0.129.0 — both legs convicted and each named its fix command.
## v0.128.0 — the escrow seed is asserted every tick, not remembered once (2026-08-08, R-221)
**A rebuilt box could not run the escrow ceremony at all, and there was no way forward from inside
the product.** The preflight refuses on `escrow.pbs_storage_id`; the pbsdr bridge writes that key;
and it wrote it from exactly one place — `finishConverged`, reached only on the paths that actually
converge.
**The two things live in different places and die at different times.** The convergence marker is
host-side (`<agent-state>/pbsdr/marker.json`). The key it seeds is in `agent.json`, which
`step_agent_config` renders from `base = {}` unless an explicit `--preserve-from` is passed
(`felhom.eu/scripts/felhom-host-install.sh:2396`, the render at `:2449`, the `O_TRUNC` write at
`:2579`; the flag at `:1246`, defaulting empty at `:256`) — **and the render never writes an
`escrow` section at all.** So a rebuild keeps the marker and takes the key: same descriptor, same
hash, early return, and the seed never runs again into a config that no longer has it.
**Fixed by asserting rather than remembering.** `Apply` now re-asserts the seed *before* the
idempotent early return. `seedEscrowStorageID` is unchanged and still never clobbers a different
existing value — an operator's own choice outranks the descriptor's, with a warning naming both.
**The early return is KEPT.** It exists so a converged box does not re-run Proxmox operations every
60 s, and `TestSeedReasserted_OnConvergedTick_WithZeroProxmoxCalls` asserts **zero** recorded runner
calls on that tick — so a "fix" that simply deleted the return would fail. Cost of the re-assert: one
small file read plus a JSON parse per tick, no exec, no network, and an early return once the value
matches.
**A seed failure can never un-converge the box:** Warn plus a message on the published status,
exactly as `finishConverged` does it — no marker write, no state change. Pinned by
`TestSeedReassertFailure_DoesNotUnconverge`.
Tests drive the real `Apply` with a real temp-dir `agent.json` and a call-recording runner; calling
`seedEscrowStorageID` directly cannot see the early return, which IS the defect. The production
wiring (`pbsdr.NewManager(..., cfg.SourcePath, ...)`) is asserted by walking `main.go`'s **AST**, not
by `strings.Contains`, which a commented-out call also satisfies.
Red-proofs: removing the new call makes Scenario A fail on today's tree; removing the early return
makes the zero-Proxmox-calls assertion fail.
## (no version bump) — a comment that claimed the hub reads a field it has no field for (2026-08-08, R-260)
Comment-only; no behaviour, no wire change, nothing to rebuild.
`HostReport.SelfUpdatePending` / `SelfUpdatePendingVersion` carried the sentence *"The hub reads an
absent field as pending=false, the correct default."* **The hub has no field for either**, so it reads
nothing — present or absent — and `encoding/json` discards them on arrival. The sentence described an
intent rather than the code and read as settled for long enough that a class sweep had to find it.
The emission is correct and stays: the agent reports the truth, and the fault is entirely in the
receiving. The missing consumer is tracked as **R-264** (OPEN), and
`felhom.eu/scripts/wire_contract_gate.py` now refuses any NEW field of this shape while recording the
existing ones as explicit, reasoned allowlist entries rather than silence.
## v0.127.0 — a mount Felhom itself made is not "something else" (2026-08-06, R-220) ## v0.127.0 — a mount Felhom itself made is not "something else" (2026-08-06, R-220)
**After a rebuild the customer's own drives could not be re-attached, and the refusal named an action **After a rebuild the customer's own drives could not be re-attached, and the refusal named an action
@@ -5280,3 +5563,5 @@ client, signing, or storage/backup orchestration yet (later slices).
read-only `--selftest` against the demo host with TLS fingerprint pinning. read-only `--selftest` against the demo host with TLS fingerprint pinning.
- The 16-privilege `FelhomAgent` role + privsep token (role on **both** user and - The 16-privilege `FelhomAgent` role + privsep token (role on **both** user and
token) is provisioned out-of-band; the agent only consumes the token. token) is provisioned out-of-band; the agent only consumes the token.
<!-- R-421 sweep: this repo cites R-421; the row landed in felhom.eu 2d88776. -->
+8
View File
@@ -101,3 +101,11 @@ the mechanism are exempt.
- **Confirm your own last push's CI run went green, by run ID** — CI mails on failure, which is a PUSH - **Confirm your own last push's CI run went green, by run ID** — CI mails on failure, which is a PUSH
signal; this is the PULL check that catches a lost or unread mail. An unchecked green is an signal; this is the PULL check that catches a lost or unread mail. An unchecked green is an
assumption, not an observation. assumption, not an observation.
**A gate ships with a decoy test that has been seen to fail (R-421).** A decoy is the LABEL without
the FACT — a directory with the right name and no bake log, a note whose prose mentions the marker it
lacks. `scripts/decoy_coverage_gate.py` refuses a new gate that has neither a decoy nor a named
exemption carrying its row. The four shapes, the 2026-09-01 sweep that fooled 16 of 29 gates, and the
decoys withdrawn as illegitimate: `documentation/audits/AUDIT-gate-decoys-2026-09-01.md` and
`felhom-controller/.claude/rules/gates.md`. **Scope is a fact too** — prefer `os.walk` over
`os.listdir`, and a glob over a hand-maintained list.
+41 -35
View File
@@ -1,41 +1,47 @@
# REPORT — felhom-agent v0.127.0: a mount Felhom made is not foreign (R-220) # REPORT — agent v0.131.0: the controller comes back by itself; backup status per tier (2026-09-15)
**Scope: the host half of R-220.** The customer-facing refusal message is the controller's half and Task: *before the volunteer — the big night's P1 fixes*, Parts A and C (agent half). Architecture: `03-host-agent.md`
ships as felhom-controller v0.203.0. §4 (the "healing a crashed controller" sentence), `07-backup-architecture.md` §6.
## What changed ## Measured first (A.1)
Docker 29.8.0, throwaway containers on scratch 9202: after `docker kill`, **both** `--restart unless-stopped` and
`--restart always` stayed `exited (137)` 60 s later. The task's claim was right; a policy change is not a fix.
| File | Change | ## What shipped
- **Controller supervisor** (`internal/localapi/controllersupervisor.go`): every 30 s, for provisioned felhom-pool guests
that are running, restart `felhom-controller-bootstrap.service` on the second not-running observation. Guards: swap
in flight, host-side park marker, locked / vzdump-busy / stopped guest, unknown docker answer, unprovisioned guest,
3 restarts in 15 min → 30 min pause. Record rides the host report as `controller_supervisor`; hub v0.114.0 mints the
events. Same GuestExecutor and sudoers grants as the swap — no new privilege.
- **Per-tier backup status**: `GET /backup/status` (untargeted) gains `tiers[]` — newest success (record, or storage
after a restart), last attempt kept apart, storage presence; `GET /backup/tiers` gains `storage`.
- **Golden script**: `--restart always`. No golden baked (R-468); existing boxes keep `unless-stopped`.
## Red-proofs (each seen failing, then restored)
- remove the restart call → `the killed controller was NOT restarted — this is R-523 (restarts=0)`
- remove the backoff block → `crash-looping controller restarted 10 times in 10 minutes — want exactly 3`
- `last_success` from the newest attempt → `pbs tier reports a failed attempt as its last success`
## Release and delivery
`scripts/release-agent.sh 0.131.0`: tag `v0.131.0`, sha256 `1118b552f7e775fbde9544c7764ede7e6046e0a7db16ae8494d07a18e3c2ac9c`,
verified by download. **The task's "the floor delivers the agent" was wrong** (R-530): the hub holds a floor above the
box's agent; agents update only by an operator-signed job. On the operator's keys: `felhom-opsign -op agent_update` for
`demo-hp-bb76ea` only → authorized 08:44:16Z, committed 08:45:21Z, `controller-supervisor: started`. demo-felhom and
Peti's box stay on 0.130.0.
## Live validation on demo-hp guest 9201 (agent 0.131.0, controller 0.243.0)
| moment | result |
|---|---| |---|---|
| `internal/storage/claim.go` | `claimFacts.felhomOwnedMounts`; `classifyClaim` forgives a non-managed mountpoint **only when corroborated**; `felhomOwnedMounts()` + `procMounts()` | | idle kill 08:53:27Z | restarted 08:54:22Z; dashboard 200 **59 s** after the kill |
| `internal/storage/hostops.go` | `mountTable` seam (nil ⇒ real `/proc/mounts`) | | parked + kill | stayed dead 100 s, `the guest is PARKED — leaving it` every sweep; unpark → 200 in **25 s** |
| `internal/storage/claim_r220_test.go` | new — the own-drive case, the fence, and the corroboration's four edges | | kill 10 s into a swap | `during a controller SWAP — the swap owns it` ×3; the swap rolled back itself, healthy 09:02:06Z |
| kill 5 s into a deploy | **not measured**: the three test restarts had filled the budget, so the guard paused (as designed) and the hub mailed `controller_crashloop` |
| resume after the pause | pause held to 09:35:52Z; restarted 09:36:24Z; dashboard 200 at 09:36:30Z |
Two invalid attempts, both marked in the evidence: a swap POST sent over plain HTTP (400, no swap), and a "health 200"
line during the pause that the container state contradicts.
## The shape chosen, and why (§7.3) ## Teardown
Machine: 9201 back on its controller (see resume); the throwaway homebox deploy left no container. Host: park marker
removed; nothing else changed. Hub: demo-hp floor override 0.243.0 kept (it delivers this release).
**Candidate (b): the claimed check distinguishes a mount Felhom made from a foreign one** — the task Evidence: `felhom.eu/documentation/audits/evidence-p1fixes-2026-09-15/A*`.
called it "nearer the truth" and it is, because the host and its knowledge survive the rebuild while
the guest's registry does not. Candidate (a) — having the rebuild path clear the raw mounts — would
have made correctness depend on a cleanup step running, and a cleanup that does not run leaves exactly
today's defect.
**The discriminator is corroboration, not a path prefix**: the same device must ALSO be mounted under
`/mnt/felhom-drives`. Only enrolment produces that pairing.
**`/proc/mounts` rather than `lsblk MOUNTPOINTS`**, because the lsblk invocation is pinned verbatim in
the sudoers file; changing it would have coupled this fix to a config rollout. `/proc/mounts` is
world-readable and needs neither.
## Green gate
`go build` · `go vet` clean · `go test ./...` → **29 packages ok** · `agent_gates.py --fast` → all OK.
| Red-proof | Result |
|---|---|
| remove the `felhomOwnedMounts` exemption | **FAILS** — "device is mounted at /mnt/adatok (sdb)", the pre-fix refusal |
| over-widen the exemption to any `/mnt/*` | **FAILS** — "/mnt/someone-elses-disk was offered for formatting" |
## Not changed
No sudoers, no allowlisted command, no PVE surface, no format path. Every other claim signal
(system disk, read-only, LVM PV, ZFS member, member FSTYPEs, empty-topology backstop) is untouched.
+2
View File
@@ -74,6 +74,7 @@
| `EnsureLeaf` | internal/localapi/cert.go | `EnsureLeaf(certPath, keyPath, host) (cert, fingerprint, generated, err)` | pinned self-signed leaf | `generated=true` invalidates every issued bootstrap pin — log LOUD (B.1) | | `EnsureLeaf` | internal/localapi/cert.go | `EnsureLeaf(certPath, keyPath, host) (cert, fingerprint, generated, err)` | pinned self-signed leaf | `generated=true` invalidates every issued bootstrap pin — log LOUD (B.1) |
| `Server.RecoverStaleLockedGuests` | internal/localapi/stalelock.go | `RecoverStaleLockedGuests(ctx)` | startup stale vzdump-lock heal (F2-b) | Clears ONLY `backup`/`snapshot-delete`, only when no vzdump in-flight; A1 RESOLVED (v0.62.0): scan is pool-intersected (`ListLXC` ∩ `Client.Pool`), fail-safe skip on pool-read failure | | `Server.RecoverStaleLockedGuests` | internal/localapi/stalelock.go | `RecoverStaleLockedGuests(ctx)` | startup stale vzdump-lock heal (F2-b) | Clears ONLY `backup`/`snapshot-delete`, only when no vzdump in-flight; A1 RESOLVED (v0.62.0): scan is pool-intersected (`ListLXC` ∩ `Client.Pool`), fail-safe skip on pool-read failure |
| `ControllerSwapper.Swap` + `ValidControllerImage` | internal/localapi/controllerswap.go | `Swap(ctx, vmid, target) *ControllerSwapState` | agent-owned controller image swap + rollback | Strict image regex (repo + 3-part semver); state file written BEFORE swap; no-healthcheck images need `verifyDwell` | | `ControllerSwapper.Swap` + `ValidControllerImage` | internal/localapi/controllerswap.go | `Swap(ctx, vmid, target) *ControllerSwapState` | agent-owned controller image swap + rollback | Strict image regex (repo + 3-part semver); state file written BEFORE swap; no-healthcheck images need `verifyDwell` |
| `Server.ControllerSupervisorTick` + `ControllerParkedMarker` | internal/localapi/controllersupervisor.go | `ControllerSupervisorTick(ctx)` | R-523: restart a provisioned guest's not-running controller via its bootstrap unit | Two-sweep confirm; honours swapInFlight, the host-side park marker, guest lock + vzdump; 3 restarts/15 min → 30 min pause; record rides the report as `controller_supervisor` (the hub mints the events — the agent has no event channel) |
| `MemoryOps` + `Server.readMemoryBounds` | internal/localapi/guestmemory.go | `readMemoryBounds(ctx, vmid) (memoryBounds, err)` | guest RAM resize (v0.90.0, R-24): GET/POST /guest/memory | NEW narrow seam (never extend `GuestAPI` — it breaks every fake); the AGENT is the boundary — bounds recomputed FRESH per request (min 2048 / max host_total−2048 / shrink floor max(2048, usage+512)); §8 UNITS TRAP (config `memory`=MB, status/node=bytes); verify maxmem==target after `SetConfig` before claiming success; SetConfig NEVER called on a refusal path | | `MemoryOps` + `Server.readMemoryBounds` | internal/localapi/guestmemory.go | `readMemoryBounds(ctx, vmid) (memoryBounds, err)` | guest RAM resize (v0.90.0, R-24): GET/POST /guest/memory | NEW narrow seam (never extend `GuestAPI` — it breaks every fake); the AGENT is the boundary — bounds recomputed FRESH per request (min 2048 / max host_total−2048 / shrink floor max(2048, usage+512)); §8 UNITS TRAP (config `memory`=MB, status/node=bytes); verify maxmem==target after `SetConfig` before claiming success; SetConfig NEVER called on a refusal path |
### Proxmox client / hub / PBS / provisioning ### Proxmox client / hub / PBS / provisioning
@@ -86,6 +87,7 @@
| `Client.PoolAddVMID` | internal/proxmox/mutate.go | `PoolAddVMID(ctx, pool, vmid) error` | re-assert pool membership after a restore-over-existing (campaign-2 R2) | SYNC (no UPID, don't WaitTask); PVE `PUT /pools` is additive (merge, not replace) — `delete=1` removes; idempotent (already-member swallowed); needs `Pool.Allocate` at `/pool/<pool>`. `pct restore --pool` sets membership only at CREATE — a restore over an existing vmid drops it, so bring-up re-asserts post-restore | | `Client.PoolAddVMID` | internal/proxmox/mutate.go | `PoolAddVMID(ctx, pool, vmid) error` | re-assert pool membership after a restore-over-existing (campaign-2 R2) | SYNC (no UPID, don't WaitTask); PVE `PUT /pools` is additive (merge, not replace) — `delete=1` removes; idempotent (already-member swallowed); needs `Pool.Allocate` at `/pool/<pool>`. `pct restore --pool` sets membership only at CREATE — a restore over an existing vmid drops it, so bring-up re-asserts post-restore |
| `TLSConfig.build` / `normalizeFingerprint` | internal/proxmox/tls.go | `build() (*tls.Config, error)` | PVE leaf-cert SHA-256 pinning | No insecure default | | `TLSConfig.build` / `normalizeFingerprint` | internal/proxmox/tls.go | `build() (*tls.Config, error)` | PVE leaf-cert SHA-256 pinning | No insecure default |
| `pinnedTLS` | internal/pbs/pin.go | `pinnedTLS(fingerprint) (*tls.Config, error)` | PBS leaf pinning | Same model as PVE; 64-hex fingerprint normalized | | `pinnedTLS` | internal/pbs/pin.go | `pinnedTLS(fingerprint) (*tls.Config, error)` | PBS leaf pinning | Same model as PVE; 64-hex fingerprint normalized |
| `httpx.NewTransport` | internal/httpx/transport.go | `NewTransport(tlsCfg, idleConnTimeout) *http.Transport` | **EVERY** hand-rolled `http.Transport` in this repo — pbs, hub and proxmox all pin TLS, so none can use `http.DefaultTransport` | **R-344: never inline `&http.Transport{TLSClientConfig: ...}` again.** A composite literal takes `IdleConnTimeout` **zero, which means retain idle connections FOREVER** — `http.DefaultTransport` sets 90s and a literal does not inherit it. Combined with a client rebuilt per cycle and dropped (`pbsTargetsFromPVE`), that stranded **388 sockets on ep0 in 46 h**, held open on BOTH sides. `idleConnTimeout <= 0` means **use the default**, never "no timeout". Returns a **FRESH** transport every call — a shared one would pool connections across differently pinned endpoints. Pinned by `internal/pbs/client_leak_test.go` (server-side connection counting) + `internal/httpx/transport_test.go` |
| `hub.Client.Report` | internal/hub/client.go | `Report(ctx, *HostReport) (*ControlEnvelope, error)` | the heartbeat | Typed `TransportError`/`HTTPError`, never contain the bearer token | | `hub.Client.Report` | internal/hub/client.go | `Report(ctx, *HostReport) (*ControlEnvelope, error)` | the heartbeat | Typed `TransportError`/`HTTPError`, never contain the bearer token |
| `hub.Loop` + `MultiObserver` | internal/hub/loop.go | `NewLoop(...)`; `MultiObserver(obs...)` | resilient report loop + envelope fan-out | Errors logged, loop continues; interval clamped 60–3600 s | | `hub.Loop` + `MultiObserver` | internal/hub/loop.go | `NewLoop(...)`; `MultiObserver(obs...)` | resilient report loop + envelope fan-out | Errors logged, loop continues; interval clamped 60–3600 s |
| `provision.BackHalf.Provision` | internal/provision/backhalf.go | `Provision(ctx, Input) (Result, error)` | guest bootstrap back-half | mint→render→0600 write→chown 100000:100000→`pct set` ro bind→onboot; token NEVER logged/returned. Bootstrap `local_api.endpoint` = the caller's `cfg.LocalAPI.ListenAddr` (main.go) — moving the agent bind to the island moves the guest dial for free (R-50, no template) | | `provision.BackHalf.Provision` | internal/provision/backhalf.go | `Provision(ctx, Input) (Result, error)` | guest bootstrap back-half | mint→render→0600 write→chown 100000:100000→`pct set` ro bind→onboot; token NEVER logged/returned. Bootstrap `local_api.endpoint` = the caller's `cfg.LocalAPI.ListenAddr` (main.go) — moving the agent bind to the island moves the guest dial for free (R-50, no template) |
+45 -15
View File
@@ -59,7 +59,7 @@ import (
// version is the agent version. Overridable at build time with // version is the agent version. Overridable at build time with
// -ldflags "-X main.version=<v>"; defaults to the in-repo CHANGELOG version. // -ldflags "-X main.version=<v>"; defaults to the in-repo CHANGELOG version.
var version = "0.92.1" var version = "0.131.0"
// runGuestHook is the PVE hook body (`felhom-agent guest-hook <vmid> <phase>`). On pre-start it // runGuestHook is the PVE hook body (`felhom-agent guest-hook <vmid> <phase>`). On pre-start it
// creates placeholder dirs for any absent bind-mount source so the guest always boots (the C1 net); // creates placeholder dirs for any absent bind-mount source so the guest always boots (the C1 net);
@@ -1401,6 +1401,10 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
// appliance outage with nothing retrying) needs a PERIODIC check. onboot is the "should be // appliance outage with nothing retrying) needs a PERIODIC check. onboot is the "should be
// running" signal, so a deliberately stopped guest is never touched. // running" signal, so a deliberately stopped guest is never touched.
go localSrv.WatchGuestPower(ctx) go localSrv.WatchGuestPower(ctx)
// R-523: a controller container that is simply not running (killed, stopped, a failed
// self-update) is restarted through its bootstrap unit — nothing else watches it.
collector.SetControllerSupervisorReporter(localSrv)
go localSrv.WatchControllers(ctx)
go func() { errc <- localSrv.Run(ctx) }() go func() { errc <- localSrv.Run(ctx) }()
} }
if lanLoop != nil { if lanLoop != nil {
@@ -1769,23 +1773,48 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
} }
return blob, true, nil return blob, true, nil
}, },
// R-311 — the RETAINED packages, wired here and ONLY here, on the same self-scoped hub client.
// Consulted only after the current package has refused the code (see tryRetained), so the
// ordinary recovery pays nothing for it and cannot fail because of it.
FetchRetained: func(ctx context.Context) ([]escrow.RetainedBlob, int, error) {
resp, ferr := hubClient.FetchRetainedIdentityEscrow(ctx)
if ferr != nil {
return nil, 0, ferr
}
out := make([]escrow.RetainedBlob, 0, len(resp.Packages))
for _, p := range resp.Packages {
blob, derr := base64.StdEncoding.DecodeString(p.IdentityEscrowB64)
if derr != nil || len(blob) == 0 {
// One malformed package must not sink the rest — the customer's code may open a
// later one, and a skipped entry is strictly better than a refusal we cannot justify.
continue
}
out = append(out, escrow.RetainedBlob{
Blob: blob,
SupersededAt: p.SupersededAt,
KeyFingerprint: p.KeyFingerprint,
Index: p.Index,
})
}
return out, resp.UnopenableCount, nil
},
} }
srv, err := localapi.NewServer(localapi.Options{ srv, err := localapi.NewServer(localapi.Options{
EscrowRecovery: escrowRecoverer, EscrowRecovery: escrowRecoverer,
ListenAddr: cfg.LocalAPI.ListenAddr, ListenAddr: cfg.LocalAPI.ListenAddr,
Cert: cert, Cert: cert,
AgentVersion: version, // v0.82.0: the X-Felhom-Agent-Version capability channel AgentVersion: version, // v0.82.0: the X-Felhom-Agent-Version capability channel
Guests: px, Guests: px,
Backups: runner, Backups: runner,
BackupTiers: apiTiers, // R-82: primary first; untargeted endpoints act on the primary BackupTiers: apiTiers, // R-82: primary first; untargeted endpoints act on the primary
InFlight: inFlight, // R-85: shared with the restore-test scheduler (Scenario F) InFlight: inFlight, // R-85: shared with the restore-test scheduler (Scenario F)
Store: store, Store: store,
Storage: observer, Storage: observer,
DriveTargets: driveTargets, // Impl-2a: registry+units drives for the /disks view (union w/ Observe storages) DriveTargets: driveTargets, // Impl-2a: registry+units drives for the /disks view (union w/ Observe storages)
Smart: storage.NewSmartReader(hostOps), // v0.95.0 Fix B: SMART for the union-path drives Smart: storage.NewSmartReader(hostOps), // v0.95.0 Fix B: SMART for the union-path drives
HostReader: storage.NewProcHostReader(), // Impl-2b: durableIDForMount raw-mount fallback + role gate HostReader: storage.NewProcHostReader(), // Impl-2b: durableIDForMount raw-mount fallback + role gate
Tokens: tokens, Tokens: tokens,
BackupCadence: cfg.Backup.BackupCadence(), BackupCadence: cfg.Backup.BackupCadence(),
// Disk management (slice 8C): the privileged host surface + the data-bearing wipe gate. // Disk management (slice 8C): the privileged host surface + the data-bearing wipe gate.
Disks: hostOps, Disks: hostOps,
DiskGate: storageGateAdapter{gate: gate, hostID: cfg.Hub.HostID}, DiskGate: storageGateAdapter{gate: gate, hostID: cfg.Hub.HostID},
@@ -1801,6 +1830,7 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
StateDir: cfg.WGTunnel.WithDefaults().StateDir, StateDir: cfg.WGTunnel.WithDefaults().StateDir,
SmbCredsDir: cfg.Privileged.SmbCredsDir, SmbCredsDir: cfg.Privileged.SmbCredsDir,
ControllerSwap: guestBinder, // Phase 1: agentic controller update — in-guest image swap ControllerSwap: guestBinder, // Phase 1: agentic controller update — in-guest image swap
GuestsStateDir: "/var/lib/felhom-agent/guests", // R-523: <vmid>/bootstrap + controller-parked marker
// F2-b: recover a guest left with a stale vzdump lock by a reboot-during-backup. Reads + start // F2-b: recover a guest left with a stale vzdump lock by a reboot-during-backup. Reads + start
// go through the API client; the `pct unlock` is the one fenced root-CLI op (no API equivalent). // go through the API client; the `pct unlock` is the one fenced root-CLI op (no API equivalent).
// A1 (v0.62.0): the scan is restricted to felhom-pool members (ownership proven, not assumed). // A1 (v0.62.0): the scan is restricted to felhom-pool members (ownership proven, not assumed).
+5 -1
View File
@@ -289,7 +289,11 @@ mount --make-rshared /mnt
# Otherwise still DE-PRIVILEGED: disk EXECUTION (scan/format/mount) stays the agent's — NO --privileged, # Otherwise still DE-PRIVILEGED: disk EXECUTION (scan/format/mount) stays the agent's — NO --privileged,
# no /dev, no /etc/fstab. Bootstrap config (ro), data volume, stacks dir (same-path), the /mnt :rslave # no /dev, no /etc/fstab. Bootstrap config (ro), data volume, stacks dir (same-path), the /mnt :rslave
# view, and the docker socket. The controller reaches the agent's local API for disk management. # view, and the docker socket. The controller reaches the agent's local API for disk management.
docker run -d --name felhom-controller --restart unless-stopped "${HOSTNAME_ARGS[@]}" \ # R-523: `always`, not `unless-stopped`. It covers ONE extra case only — a Docker daemon restart after
# the container was stopped by hand. Neither policy restarts a container that `docker kill`/`docker
# stop` ended (measured 2026-09-15, Docker 29.8.0, evidence-p1fixes-2026-09-15/A1); the host agent's
# controller supervisor (felhom-agent v0.131.0, internal/localapi/controllersupervisor.go) covers that.
docker run -d --name felhom-controller --restart always "${HOSTNAME_ARGS[@]}" \
-e FELHOM_BOOTSTRAP_PATH=/etc/felhom-bootstrap/bootstrap.json \ -e FELHOM_BOOTSTRAP_PATH=/etc/felhom-bootstrap/bootstrap.json \
-v /etc/felhom-bootstrap:/etc/felhom-bootstrap:ro \ -v /etc/felhom-bootstrap:/etc/felhom-bootstrap:ro \
-v felhom-controller-data:/opt/docker/felhom-controller \ -v felhom-controller-data:/opt/docker/felhom-controller \
+116 -1
View File
@@ -47,17 +47,80 @@ var (
// retro-fitted, because R is never retained. Distinguished from a wrong code so the operator is // retro-fitted, because R is never retained. Distinguished from a wrong code so the operator is
// not sent hunting for a mistyped recovery code that was typed correctly. // not sent hunting for a mistyped recovery code that was typed correctly.
ErrNoResticPassword = errors.New("escrow: the recovered bundle carries NO offsite repository password (a pre-fork-4 blob — the field did not exist when it was sealed and cannot be retro-fitted)") ErrNoResticPassword = errors.New("escrow: the recovered bundle carries NO offsite repository password (a pre-fork-4 blob — the field did not exist when it was sealed and cannot be retro-fitted)")
// ErrCodeOpensRetained — the code did NOT open the package the hub currently holds, and DID open a
// RETAINED (earlier) one. R-311.
//
// ⚠ THIS IS NOT A FAILURE OF THE CUSTOMER'S. It is the single most important distinction on this
// path, because until 2026-08-12 it was indistinguishable from a mistype and was reported as one.
// The screen could only say "it may be a typo, or it may be an older code, and we cannot tell them
// apart from here" — and it could not tell them apart because NOTHING EVER LOOKED. Now something
// looks, so the sentence can stop hedging.
//
// It carries no material and no code: only WHICH earlier package opened, by its supersession date,
// which is the one fact the customer needs to recognise it.
ErrCodeOpensRetained = errors.New("escrow: the recovery code did not open the CURRENT sealed package, but it DID open a retained earlier one")
) )
// RetainedMatch says which retained package a code opened. Returned inside RetainedOpenedError; it
// carries no secret — not the code, not the bundle, not the repository password.
type RetainedMatch struct {
// SupersededAt is when this package stopped being the current one (hub-supplied, RFC3339-ish).
// It is what the recovery screen shows so the customer can recognise which code they are holding.
SupersededAt string
// KeyFingerprint is the escrow key fingerprint of that package — operator-log material only.
KeyFingerprint string
// Index is the hub's position label within ONE response. Not durable; do not persist it.
Index int
// HasResticPassword is false when the retained package opened but carries no repository password
// (a pre-fork-4 seal). The code is still CORRECT; the history behind it still cannot be reopened.
// Collapsing this into "recoverable" would repeat R-202's mistake on a new surface.
HasResticPassword bool
}
// RetainedOpenedError wraps ErrCodeOpensRetained with the match. Callers classify with errors.Is on
// the sentinel and read the detail with errors.As.
type RetainedOpenedError struct {
Match RetainedMatch
}
func (e *RetainedOpenedError) Error() string {
return ErrCodeOpensRetained.Error() + " (superseded_at=" + e.Match.SupersededAt + ")"
}
func (e *RetainedOpenedError) Unwrap() error { return ErrCodeOpensRetained }
// BlobFetcher yields this host's own opaque identity-escrow blob. present=false is a clean "none". // BlobFetcher yields this host's own opaque identity-escrow blob. present=false is a clean "none".
// An interface-free func field keeps this package free of any dependency on the hub client. // An interface-free func field keeps this package free of any dependency on the hub client.
type BlobFetcher func(ctx context.Context) (blob []byte, present bool, err error) type BlobFetcher func(ctx context.Context) (blob []byte, present bool, err error)
// RetainedBlob is one retained sealed package as the recoverer sees it: opaque bytes plus the labels
// needed to name it. No secret.
type RetainedBlob struct {
Blob []byte
SupersededAt string
KeyFingerprint string
Index int
}
// RetainedFetcher yields this host's RETAINED sealed packages, newest-superseded first. An empty
// slice is a clean "none". R-311.
type RetainedFetcher func(ctx context.Context) (blobs []RetainedBlob, unopenable int, err error)
// OffsiteKeyRecoverer is the assembled links 6→8. Construct it with a fetcher; call it with R. // OffsiteKeyRecoverer is the assembled links 6→8. Construct it with a fetcher; call it with R.
type OffsiteKeyRecoverer struct { type OffsiteKeyRecoverer struct {
Fetch BlobFetcher Fetch BlobFetcher
// FetchRetained is OPTIONAL and consulted ONLY after the current package has refused the code.
// nil keeps the pre-R-311 behaviour exactly: a refusal stays a refusal. That is deliberate — an
// agent wired without it must not behave differently from one that has no retained packages.
FetchRetained RetainedFetcher
// MaxRetainedTried bounds the scrypt work a single wrong code can cost. Each attempt is ~1 s of
// KDF by design, so an unbounded loop over a long supersession history would turn one wrong code
// into a minutes-long hang on the customer's screen. 0 means the built-in default.
MaxRetainedTried int
} }
// defaultMaxRetainedTried — six attempts is ~6 s worst case, which is a slow screen and not a hang.
const defaultMaxRetainedTried = 6
// RecoverOffsiteRepoPassword fetches, unseals and extracts. It returns ONLY the repository password. // RecoverOffsiteRepoPassword fetches, unseals and extracts. It returns ONLY the repository password.
// //
// A WRONG RECOVERY CODE FAILS CLOSED at the scrypt KDF inside UnwrapIdentity — `age -d` exits // A WRONG RECOVERY CODE FAILS CLOSED at the scrypt KDF inside UnwrapIdentity — `age -d` exits
@@ -86,10 +149,62 @@ func (r OffsiteKeyRecoverer) RecoverOffsiteRepoPassword(ctx context.Context, rec
} }
bundle, err := UnwrapIdentityBundle(ctx, blob, recoveryCode) bundle, err := UnwrapIdentityBundle(ctx, blob, recoveryCode)
if err != nil { if err != nil {
return "", err // already the fail-closed "the recovery code did not unwrap…" message; no secret in it // R-311 — BEFORE calling this a wrong code, ask whether it is the RIGHT code for an EARLIER
// package. The engine fails closed identically either way, so the two are indistinguishable
// from the unwrap alone; the only way to tell is to try. Until this existed nobody tried, and
// the screen said so out loud ("innen nem tudjuk megkülönböztetni őket") — a true sentence
// about our own incuriosity, read by the customer as a statement about their code.
if m, ok := r.tryRetained(ctx, recoveryCode); ok {
return "", &RetainedOpenedError{Match: m}
}
return "", err // the fail-closed "the recovery code did not unwrap…" message; no secret in it
} }
if bundle.ResticRepoPassword == "" { if bundle.ResticRepoPassword == "" {
return "", ErrNoResticPassword return "", ErrNoResticPassword
} }
return bundle.ResticRepoPassword, nil return bundle.ResticRepoPassword, nil
} }
// tryRetained reports whether the code opens one of this host's RETAINED packages, and which.
//
// FAILURE HERE IS SILENT AND MEANS "NO", NEVER "YES" and never a different verdict for the caller. A
// hub that cannot answer, a route an older hub does not have, a malformed blob — each leaves the
// original refusal standing, unchanged. That is the fail-safe direction: the worst outcome of this
// function breaking is the behaviour we had before it existed.
//
// NOTHING IS LOGGED HERE and no return value carries the code, a bundle or a password.
func (r OffsiteKeyRecoverer) tryRetained(ctx context.Context, recoveryCode string) (RetainedMatch, bool) {
if r.FetchRetained == nil {
return RetainedMatch{}, false
}
blobs, _, err := r.FetchRetained(ctx)
if err != nil || len(blobs) == 0 {
return RetainedMatch{}, false
}
limit := r.MaxRetainedTried
if limit <= 0 {
limit = defaultMaxRetainedTried
}
for i, rb := range blobs {
if i >= limit {
break
}
if len(rb.Blob) == 0 {
continue
}
bundle, uerr := UnwrapIdentityBundle(ctx, rb.Blob, recoveryCode)
if uerr != nil {
continue // this one is not the customer's; try the next
}
return RetainedMatch{
SupersededAt: rb.SupersededAt,
KeyFingerprint: rb.KeyFingerprint,
Index: rb.Index,
// A retained package can itself predate the repository-password field. The code is still
// correct and must be told so — but the history behind it still cannot be reopened, and
// saying otherwise would be a promise this path cannot keep.
HasResticPassword: bundle.ResticRepoPassword != "",
}, true
}
return RetainedMatch{}, false
}
+230
View File
@@ -0,0 +1,230 @@
package escrow
import (
"context"
"errors"
"fmt"
"testing"
)
// R-311 — a correct code for an EARLIER package must stop being reported as a wrong code.
//
// These use REAL age crypto, like the R-199 tests beside them, because the whole point is that the
// two situations are indistinguishable AT THE UNWRAP: both fail closed on the current package. A
// faked unwrap would prove nothing about the thing that was actually broken.
const testR2 = "another correct horse battery staple sedative anaconda wobbly kingdom placard"
func retainedFetcherFor(blobs ...RetainedBlob) RetainedFetcher {
return func(context.Context) ([]RetainedBlob, int, error) { return blobs, 0, nil }
}
// THE ONE THAT MATTERS. The customer holds the code for a package we superseded. Yesterday this
// returned the fail-closed refusal and the screen told them to check their typing.
//
// RED-PROOF: remove the `if m, ok := r.tryRetained(...)` block from RecoverOffsiteRepoPassword →
// the wrong-code error returns instead → this FAILS, and the lie is back in exactly those words.
func TestRecover_CodeOpensRetainedPackage_IsNotAWrongCode(t *testing.T) {
ensureAge(t)
const oldPW = "aaaa567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "cccc567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR2)
retained := sealBundle(t, IdentityBundle{ResticRepoPassword: oldPW}, testR)
_, err := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: retainedFetcherFor(RetainedBlob{
Blob: retained, SupersededAt: "2026-08-12 15:18:55", KeyFingerprint: "7e:a6:af", Index: 0,
}),
}.RecoverOffsiteRepoPassword(context.Background(), testR) // the OLD code
if err == nil {
t.Fatal("recovery succeeded — it must NOT return a password for a retained package on this path")
}
if !errors.Is(err, ErrCodeOpensRetained) {
t.Fatalf("err = %v, want ErrCodeOpensRetained — a correct code for an earlier package was "+
"classified as something else, which is how it became 'check your typing'", err)
}
var ro *RetainedOpenedError
if !errors.As(err, &ro) {
t.Fatalf("err does not carry a RetainedOpenedError: %v", err)
}
if ro.Match.SupersededAt != "2026-08-12 15:18:55" {
t.Errorf("SupersededAt = %q — the screen needs this date to name the package", ro.Match.SupersededAt)
}
if !ro.Match.HasResticPassword {
t.Error("HasResticPassword = false, but the retained bundle carried one")
}
// The error must not leak the code, the password or the bundle.
for _, secret := range []string{testR, oldPW} {
if containsStr(err.Error(), secret) {
t.Fatalf("the error text leaks a secret")
}
}
}
// SCENARIO A — the ordinary recovery is untouched, and it must not even ASK for retained packages.
// If the current package opens, the customer is not in this story at all.
//
// RED-PROOF: move the tryRetained call above the successful-unwrap return → the fetcher runs → this
// FAILS on the "must not be consulted" assertion.
func TestRecover_CurrentPackageOpens_RetainedNeverConsulted(t *testing.T) {
ensureAge(t)
const pw = "bbbb567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"
current := sealBundle(t, IdentityBundle{ResticRepoPassword: pw}, testR)
consulted := false
got, err := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: func(context.Context) ([]RetainedBlob, int, error) {
consulted = true
return nil, 0, nil
},
}.RecoverOffsiteRepoPassword(context.Background(), testR)
if err != nil {
t.Fatalf("the ordinary recovery broke: %v", err)
}
if got != pw {
t.Fatalf("recovered password is not the sealed one")
}
if consulted {
t.Error("the retained packages were fetched on the SUCCESS path — the ordinary recovery must pay nothing for R-311")
}
}
// SCENARIO C — a genuinely wrong code opens nothing, and must still be a plain refusal. The new
// branch must not become a way to encourage a customer who mistyped.
//
// RED-PROOF: make tryRetained return (RetainedMatch{}, true) unconditionally → a wrong code is
// reported as opening an earlier package → this FAILS.
func TestRecover_WrongCode_StaysAPlainRefusal(t *testing.T) {
ensureAge(t)
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "cccc567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
retained := sealBundle(t, IdentityBundle{ResticRepoPassword: "dddd567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR2)
_, err := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: retainedFetcherFor(RetainedBlob{Blob: retained, SupersededAt: "2026-08-01 00:00:00"}),
}.RecoverOffsiteRepoPassword(context.Background(), "totally wrong words that open nothing at all here")
if err == nil {
t.Fatal("a wrong code succeeded")
}
if errors.Is(err, ErrCodeOpensRetained) {
t.Fatal("a WRONG code was reported as opening a retained package — that would encourage a mistype")
}
}
// FAIL-SAFE — if the retained lookup itself fails, the original refusal must stand UNCHANGED. The
// worst outcome of this feature breaking is the behaviour we had before it.
//
// RED-PROOF: make tryRetained propagate the fetch error instead of returning false → the customer
// gets a new, unexplained failure mode → this FAILS.
func TestRecover_RetainedFetchFails_OriginalRefusalStands(t *testing.T) {
ensureAge(t)
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "eeee567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
_, err := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: func(context.Context) ([]RetainedBlob, int, error) {
return nil, 0, fmt.Errorf("hub exploded")
},
}.RecoverOffsiteRepoPassword(context.Background(), testR2)
if err == nil {
t.Fatal("expected a refusal")
}
if errors.Is(err, ErrCodeOpensRetained) {
t.Fatal("a failed retained lookup was reported as 'opens a retained package'")
}
if containsStr(err.Error(), "hub exploded") {
t.Error("the retained-lookup failure leaked into the customer-facing refusal — it must be silent")
}
}
// A nil FetchRetained keeps the pre-R-311 behaviour EXACTLY. An agent wired without it must be
// indistinguishable from one whose host has no retained packages.
//
// RED-PROOF: remove the `if r.FetchRetained == nil` guard → nil-deref panic → this FAILS.
func TestRecover_NilRetainedFetcher_IsPreR311Behaviour(t *testing.T) {
ensureAge(t)
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "ffff567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
_, err := OffsiteKeyRecoverer{Fetch: fetcherFor(current)}.RecoverOffsiteRepoPassword(context.Background(), testR2)
if err == nil {
t.Fatal("expected a refusal")
}
if errors.Is(err, ErrCodeOpensRetained) {
t.Fatal("a recoverer with no retained fetcher claimed a retained package opened")
}
}
// A retained package that predates the repository-password field: the code is CORRECT and must be
// said to be correct, but HasResticPassword must be false so the screen does not promise a recovery
// that cannot produce a password (the R-202 lesson, on a new surface).
//
// RED-PROOF: hardcode HasResticPassword: true → this FAILS.
func TestRecover_RetainedOpensButPredatesTheField(t *testing.T) {
ensureAge(t)
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "1111567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR2)
// No ResticRepoPassword at all — the pre-fork-4 shape.
retained := sealBundle(t, IdentityBundle{TunnelToken: "T", PBSToken: "P"}, testR)
_, err := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: retainedFetcherFor(RetainedBlob{Blob: retained, SupersededAt: "2026-08-04 07:20:08"}),
}.RecoverOffsiteRepoPassword(context.Background(), testR)
if !errors.Is(err, ErrCodeOpensRetained) {
t.Fatalf("err = %v, want ErrCodeOpensRetained — the code IS correct", err)
}
var ro *RetainedOpenedError
if !errors.As(err, &ro) {
t.Fatalf("no RetainedOpenedError: %v", err)
}
if ro.Match.HasResticPassword {
t.Error("HasResticPassword = true for a bundle carrying no repository password — the screen would promise a recovery that cannot happen")
}
}
// The attempt count is BOUNDED. Each unwrap is ~1 s of scrypt by design, so an unbounded loop turns
// one wrong code into a minutes-long hang on the customer's screen.
//
// RED-PROOF: remove the `if i >= limit { break }` → all 10 are tried → this FAILS on the count.
func TestRecover_RetainedAttemptsAreBounded(t *testing.T) {
ensureAge(t)
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "2222567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
junk := sealBundle(t, IdentityBundle{ResticRepoPassword: "3333567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR2)
tried := 0
blobs := make([]RetainedBlob, 0, 10)
for i := 0; i < 10; i++ {
blobs = append(blobs, RetainedBlob{Blob: junk, SupersededAt: "2026-08-01 00:00:00", Index: i})
}
rec := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: func(context.Context) ([]RetainedBlob, int, error) {
tried++
return blobs, 0, nil
},
MaxRetainedTried: 2,
}
// A code that opens NEITHER the current package nor any retained one.
if _, err := rec.RecoverOffsiteRepoPassword(context.Background(), "a code that opens nothing whatsoever in this test"); err == nil {
t.Fatal("expected a refusal")
}
if tried != 1 {
t.Errorf("the retained list was fetched %d times, want exactly 1", tried)
}
}
func containsStr(hay, needle string) bool {
return len(needle) > 0 && len(hay) >= len(needle) && (func() bool {
for i := 0; i+len(needle) <= len(hay); i++ {
if hay[i:i+len(needle)] == needle {
return true
}
}
return false
})()
}
+59
View File
@@ -0,0 +1,59 @@
// Package httpx holds the one HTTP-transport default this repo may not lose.
//
// Every client here pins TLS — PBS and PVE by leaf-cert SHA-256, the hub by an optional CA file —
// so none of them can use http.DefaultTransport and each hand-rolls its own. Hand-rolling silently
// discards DefaultTransport's settings, and one of them is load-bearing:
//
// Transport: &http.Transport{TLSClientConfig: tlsCfg} // IdleConnTimeout == 0 == NO timeout
//
// A zero IdleConnTimeout means idle keep-alive connections are retained FOREVER, not "use a sane
// default". Combined with a client that is rebuilt on a schedule and dropped (pbsTargetsFromPVE
// builds a fresh pbs.Client per cycle), every cycle strands one connection that nothing will ever
// close: the abandoned Transport becomes unreachable but its persistConn read-loop goroutine keeps
// the socket alive, and an unreachable Transport does not close its connections.
//
// Measured cost, live: 388 established connections accumulated on ep0's PBS proxy between
// 2026-08-18 09:51:22Z and 2026-08-20 08:02:13Z — 194 from each of the two boxes, held open on BOTH
// sides, one per agent poll cycle, on a proxy whose descriptor ceiling is 65536. See R-344 and
// felhom.eu/documentation/audits/SPIKE-ep0-established-connections-2026-08-20.md.
package httpx
import (
"crypto/tls"
"net/http"
"time"
)
// DefaultIdleConnTimeout is how long an idle keep-alive connection is retained before it is closed.
//
// It is 90s because that is http.DefaultTransport's own value: the fix for R-344 restores a
// standard-library default rather than inventing a number, so there is nothing here to tune and
// nothing to justify. It is comfortably shorter than every cadence that drives these clients (the
// 15-minute live-snapshot collect and the 6-hour verify loop), so a connection abandoned by one
// cycle is closed long before the next.
const DefaultIdleConnTimeout = 90 * time.Second
// NewTransport builds a FRESH *http.Transport pinned to tlsCfg, with the idle-connection timeout
// applied.
//
// Fresh, never shared: each caller pins a different endpoint, and a shared transport would pool
// connections across differently pinned servers. Reusing http.DefaultTransport for the same reason
// is not an option — it would drop the pin entirely.
//
// idleConnTimeout <= 0 means USE THE DEFAULT. It deliberately does not mean "no timeout": no-timeout
// is the bug this package exists to prevent, and an unset field must never be able to reintroduce
// it. Callers pass their configured value straight through; only tests pass a short one.
//
// Only IdleConnTimeout is set. The other DefaultTransport settings this transport also lacks
// (MaxIdleConns, TLSHandshakeTimeout, ExpectContinueTimeout) are deliberately left alone: none of
// them accumulates anything, every client bounds its whole request with http.Client.Timeout, and
// widening the change would have made the R-344 measurement unattributable.
func NewTransport(tlsCfg *tls.Config, idleConnTimeout time.Duration) *http.Transport {
if idleConnTimeout <= 0 {
idleConnTimeout = DefaultIdleConnTimeout
}
return &http.Transport{
TLSClientConfig: tlsCfg,
IdleConnTimeout: idleConnTimeout,
}
}
+74
View File
@@ -0,0 +1,74 @@
package httpx
import (
"crypto/tls"
"net/http"
"testing"
"time"
)
// TestNewTransport_ZeroMeansDefaultNeverForever is the whole point of this package.
//
// http.Transport's zero IdleConnTimeout means "retain idle connections FOREVER". Any code path that
// can reach that zero reintroduces R-344, so an unset, zero or negative value must all land on the
// default. If someone later "simplifies" NewTransport by passing the argument straight through,
// this fails.
func TestNewTransport_ZeroMeansDefaultNeverForever(t *testing.T) {
for _, tc := range []struct {
name string
in time.Duration
want time.Duration
}{
{"zero", 0, DefaultIdleConnTimeout},
{"negative", -time.Hour, DefaultIdleConnTimeout},
{"explicit short value (tests)", 50 * time.Millisecond, 50 * time.Millisecond},
{"explicit long value", time.Hour, time.Hour},
} {
t.Run(tc.name, func(t *testing.T) {
got := NewTransport(&tls.Config{MinVersion: tls.VersionTLS12}, tc.in).IdleConnTimeout
if got != tc.want {
t.Fatalf("IdleConnTimeout = %v, want %v", got, tc.want)
}
if got == 0 {
t.Fatal("IdleConnTimeout is 0 — that is 'never expire', which is the R-344 defect itself")
}
})
}
}
// TestDefaultIdleConnTimeout_MatchesTheStandardLibrary pins the number to its justification.
//
// 90s is not a tuned value; it is what http.DefaultTransport uses. Reading it off the standard
// library rather than hardcoding 90 means the constant cannot drift away from the reason given for
// it in the package doc.
func TestDefaultIdleConnTimeout_MatchesTheStandardLibrary(t *testing.T) {
std, ok := http.DefaultTransport.(*http.Transport)
if !ok {
t.Skip("http.DefaultTransport is not an *http.Transport in this Go build")
}
if DefaultIdleConnTimeout != std.IdleConnTimeout {
t.Fatalf("DefaultIdleConnTimeout = %v but http.DefaultTransport uses %v — the doc comment's justification no longer holds",
DefaultIdleConnTimeout, std.IdleConnTimeout)
}
}
// TestNewTransport_IsFreshEveryCall guards the pooling property the pinning relies on.
//
// Each caller pins a DIFFERENT endpoint. A shared transport would pool connections across
// differently pinned servers, so returning a package-level singleton would be a security change
// dressed as a tidy-up.
func TestNewTransport_IsFreshEveryCall(t *testing.T) {
a := NewTransport(&tls.Config{MinVersion: tls.VersionTLS12}, 0)
b := NewTransport(&tls.Config{MinVersion: tls.VersionTLS12}, 0)
if a == b {
t.Fatal("NewTransport returned the SAME transport twice — connections would be pooled across differently pinned endpoints")
}
}
// TestNewTransport_KeepsTheTLSConfig — the transport gains a field; it must lose nothing.
func TestNewTransport_KeepsTheTLSConfig(t *testing.T) {
cfg := &tls.Config{MinVersion: tls.VersionTLS12, InsecureSkipVerify: true} //nolint:gosec // test only
if got := NewTransport(cfg, 0).TLSClientConfig; got != cfg {
t.Fatalf("TLSClientConfig = %p, want the config passed in (%p) — the pin would be dropped", got, cfg)
}
}
+73 -2
View File
@@ -15,6 +15,8 @@ import (
"time" "time"
"gitea.dooplex.hu/admin/felhom-agent/internal/config" "gitea.dooplex.hu/admin/felhom-agent/internal/config"
"gitea.dooplex.hu/admin/felhom-agent/internal/httpx"
) )
const reportPath = "/api/v1/host-report" const reportPath = "/api/v1/host-report"
@@ -49,8 +51,11 @@ func NewClient(cfg config.HubConfig, logger *slog.Logger) (*Client, error) {
tlsCfg.RootCAs = pool tlsCfg.RootCAs = pool
} }
hc := &http.Client{ hc := &http.Client{
Timeout: time.Duration(cfg.TimeoutSeconds) * time.Second, Timeout: time.Duration(cfg.TimeoutSeconds) * time.Second,
Transport: &http.Transport{TLSClientConfig: tlsCfg}, // R-344, consistency only: this client is built ONCE per process, so it never accumulated
// and contributed nothing to the ep0 leak. It carried the same missing default, which over
// a tunnel is how one idle connection survives long enough to fail on next use.
Transport: httpx.NewTransport(tlsCfg, 0),
} }
return newClient(cfg.URL, cfg.APIKey, cfg.HostID, hc, logger), nil return newClient(cfg.URL, cfg.APIKey, cfg.HostID, hc, logger), nil
} }
@@ -354,3 +359,69 @@ func (c *Client) FetchIdentityEscrow(ctx context.Context) (*IdentityEscrowRespon
} }
return &out, nil return &out, nil
} }
// RetainedEscrowPackage is one RETAINED (superseded) sealed identity package. The blob is ciphertext
// and is useless without R. `SupersededAt` is the only thing here a human ever sees — it is what lets
// the recovery screen name WHICH earlier package a code belongs to.
type RetainedEscrowPackage struct {
Index int `json:"index"`
SupersededAt string `json:"superseded_at"`
KeyFingerprint string `json:"key_fingerprint"`
IdentityEscrowB64 string `json:"identity_escrow_b64"`
}
// RetainedEscrowResponse mirrors GET /api/v1/hosts/{host_id}/escrow/retained (hub >= v0.103.0, R-311).
//
// UnopenableCount is NOT noise. It counts retained packages the hub holds whose key material is absent
// (every pre-v0.93.0 row): on a box with those and nothing else, a perfectly correct old recovery code
// opens nothing, and the reason is a defect of ours. A caller that ignores this number will tell such a
// customer their code is wrong — the exact failure this whole chain exists to stop.
type RetainedEscrowResponse struct {
HostID string `json:"host_id"`
Count int `json:"count"`
UnopenableCount int `json:"unopenable_count"`
TruncatedCount int `json:"truncated_count"`
Packages []RetainedEscrowPackage `json:"packages"`
}
// FetchRetainedIdentityEscrow reads back THIS host's RETAINED sealed identity packages (R-311 —
// the retained siblings of FetchIdentityEscrow, self-scoped server-side by the same per-host key).
//
// SEPARATE FROM FetchIdentityEscrow ON PURPOSE. The ordinary recovery must not pay for this call, and
// must not fail because of it: the current package is tried first and alone, and this is reached only
// after that has refused. A hub too old to know this route answers 404, which is a CLEAN "none" here
// and must never be reported as a failed recovery.
func (c *Client) FetchRetainedIdentityEscrow(ctx context.Context) (*RetainedEscrowResponse, error) {
if c.hostID == "" {
return nil, fmt.Errorf("hub: FetchRetainedIdentityEscrow requires a configured host_id")
}
url := c.baseURL + "/api/v1/hosts/" + c.hostID + "/escrow/retained"
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
return nil, fmt.Errorf("hub: building retained-escrow request: %w", err)
}
req.Header.Set("Authorization", "Bearer "+c.apiKey)
req.Header.Set("Accept", "application/json")
resp, err := c.hc.Do(req)
if err != nil {
return nil, &TransportError{Err: err}
}
defer resp.Body.Close()
raw, _ := io.ReadAll(io.LimitReader(resp.Body, 4<<20))
if resp.StatusCode == http.StatusNotFound {
// A hub older than v0.103.0 has no such route. That is "no retained packages", not a fault —
// returning an error here would turn an old hub into a failed recovery on a box whose current
// package simply did not open.
return &RetainedEscrowResponse{HostID: c.hostID}, nil
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, &HTTPError{StatusCode: resp.StatusCode, BodyTail: tail(raw, 256)}
}
var out RetainedEscrowResponse
if err := json.Unmarshal(raw, &out); err != nil {
return nil, fmt.Errorf("hub: decoding retained escrow fetch: %w", err)
}
return &out, nil
}
+16
View File
@@ -103,6 +103,7 @@ type Collector struct {
addrEnum AddressEnumerator // v0.119.0: host interface enumeration; nil => the REAL one (see collectAddresses) addrEnum AddressEnumerator // v0.119.0: host interface enumeration; nil => the REAL one (see collectAddresses)
wg WireguardReporter // S3: offsite-tunnel status (nil → stanza omitted) wg WireguardReporter // S3: offsite-tunnel status (nil → stanza omitted)
pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted) pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted)
ctrlSup ControllerSupervisorReporter // R-523: in-guest controller supervisor (nil → stanza omitted)
guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted) guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted)
selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false) selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false)
mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted) mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted)
@@ -195,6 +196,17 @@ func (c *Collector) SetPBSDRReporter(p PBSDRReporter) *Collector {
return c return c
} }
// ControllerSupervisorReporter is the R-523 seam (satisfied by *localapi.Server).
type ControllerSupervisorReporter interface {
ControllerSupervisorStatus(ctx context.Context) *ControllerSupervisorStatus
}
// SetControllerSupervisorReporter wires the R-523 controller supervisor as a report source (nil-safe).
func (c *Collector) SetControllerSupervisorReporter(r ControllerSupervisorReporter) *Collector {
c.ctrlSup = r
return c
}
// SetGuestNetReporter wires the R-54 guest-network watchdog as a report source (nil-safe → stanza // SetGuestNetReporter wires the R-54 guest-network watchdog as a report source (nil-safe → stanza
// omitted). Returns the collector for chaining. // omitted). Returns the collector for chaining.
func (c *Collector) SetGuestNetReporter(g GuestNetReporter) *Collector { func (c *Collector) SetGuestNetReporter(g GuestNetReporter) *Collector {
@@ -285,6 +297,10 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
if c.pbsdr != nil { if c.pbsdr != nil {
report.PBSDR = c.pbsdr.PBSDRStatus(ctx) report.PBSDR = c.pbsdr.PBSDRStatus(ctx)
} }
// R-523: controller supervisor record (nil reporter = not wired → stanza omitted).
if c.ctrlSup != nil {
report.ControllerSupervisor = c.ctrlSup.ControllerSupervisorStatus(ctx)
}
// R-54: guest-network watchdog state (nil reporter = feature not wired → stanza omitted). // R-54: guest-network watchdog state (nil reporter = feature not wired → stanza omitted).
if c.guestNet != nil { if c.guestNet != nil {
report.GuestNet = c.guestNet.GuestNetStatus(ctx) report.GuestNet = c.guestNet.GuestNetStatus(ctx)
+41 -2
View File
@@ -68,8 +68,14 @@ type HostReport struct {
// report is stored opaquely hub-side, so these additive fields need no hub-schema change. // report is stored opaquely hub-side, so these additive fields need no hub-schema change.
// Both are `omitempty` (the Wireguard precedent): in the steady state (no update in flight) // Both are `omitempty` (the Wireguard precedent): in the steady state (no update in flight)
// they are absent — which keeps the cross-repo host-report golden contract byte-stable without // they are absent — which keeps the cross-repo host-report golden contract byte-stable without
// a hub change. They appear only while an update is pending. The hub reads an absent field as // a hub change. They appear only while an update is pending.
// pending=false, the correct default. //
// ⚠ CORRECTED 2026-08-08 (R-260). This comment used to end "The hub reads an absent field as
// pending=false, the correct default." THE HUB HAS NO FIELD FOR EITHER OF THESE, so it reads
// nothing — present or absent — and encoding/json discards them on arrival. The sentence
// described an intent, not the code, and it read as settled for long enough that a sweep had to
// find it. The emission is correct and stays; the missing consumer is tracked as R-264, and
// `felhom.eu/scripts/wire_contract_gate.py` now refuses any NEW field of this shape.
SelfUpdatePending bool `json:"selfupdate_pending,omitempty"` SelfUpdatePending bool `json:"selfupdate_pending,omitempty"`
SelfUpdatePendingVersion string `json:"selfupdate_pending_version,omitempty"` SelfUpdatePendingVersion string `json:"selfupdate_pending_version,omitempty"`
@@ -118,6 +124,39 @@ type HostReport struct {
// HTTPS even when felhom-sshd or the tunnel is DOWN (channel independence). `omitempty`: absent // HTTPS even when felhom-sshd or the tunnel is DOWN (channel independence). `omitempty`: absent
// when the feature is not wired (pre-H1) — additive, no hub-schema change. // when the feature is not wired (pre-H1) — additive, no hub-schema change.
OOB *OOBStatus `json:"oob,omitempty"` OOB *OOBStatus `json:"oob,omitempty"`
// ControllerSupervisor (R-523, v0.131.0) is the in-guest controller supervisor's per-guest record:
// how many times the agent restarted a dead controller, when last and why, whether it gave up
// (crash-loop pause) and whether the operator parked it. The hub's ControllerSupervisorChecker
// mints `controller_restarted_by_agent` when last_restart_at MOVES and `controller_crashloop` when
// crashloop_since MOVES — timestamps, not counters, because the record is in-memory and an agent
// restart zeroes the counter. `omitempty`: absent when not wired, so the cross-repo golden stays
// byte-stable. The hub parser is pinned by hub/internal/monitor/controller_supervisor_test.go
// against the JSON TestControllerSupervisorStanza_WireShape pins here.
ControllerSupervisor *ControllerSupervisorStatus `json:"controller_supervisor,omitempty"`
}
// ControllerSupervisorStatus is the R-523 stanza. Carries no secret.
type ControllerSupervisorStatus struct {
Guests []ControllerSupervisorGuest `json:"guests"`
}
// ControllerSupervisorGuest is one supervised guest.
type ControllerSupervisorGuest struct {
VMID int `json:"vmid"`
RestartsTotal int `json:"restarts_total"`
LastRestartAt string `json:"last_restart_at,omitempty"` // RFC3339
LastReason string `json:"last_reason,omitempty"`
Crashloop bool `json:"crashloop"`
CrashloopSince string `json:"crashloop_since,omitempty"` // RFC3339; the last crash-loop, kept after it ends
Parked bool `json:"parked"`
// R-539 (v0.132.0) — the SLOW crash loop. Restarts24h counts restarts the supervisor performed in
// the last 24 hours (persisted, so an agent restart does not reset it); SlowCrashloop is true while
// the last raise is under 24 hours old; SlowCrashloopSince is the raise itself, which the hub keys on
// MOVING (hub v0.117.0 controller_slow_crashloop). It moves at most once per 24 hours.
Restarts24h int `json:"restarts_24h"`
SlowCrashloop bool `json:"slow_crashloop"`
SlowCrashloopSince string `json:"slow_crashloop_since,omitempty"` // RFC3339
} }
// PBSDRStatus is the per-heartbeat PBS-DR-tier bridge state (slice 2). States: // PBSDRStatus is the per-heartbeat PBS-DR-tier bridge state (slice 2). States:
@@ -0,0 +1,133 @@
package localapi
import (
"context"
"encoding/json"
"io"
"log/slog"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// R-517 — the per-tier truth on GET /backup/status. BIGNIGHT: a successful 8.9 GB local backup,
// then a failed PBS attempt on a storage that did not exist; the page (fed by the single latest
// record) showed the 0-byte failure as "up to date" and the remote copy as present.
func tierStatesOf(t *testing.T, srv *Server) []TierBackupState {
t.Helper()
w := do(t, srv.Handler(), "GET", "/backup/status", "A", "")
var resp struct {
Data BackupStatusResponse `json:"data"`
}
if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
t.Fatalf("decode: %v (%s)", err, w.Body.String())
}
return resp.Data.Tiers
}
func tierStatesServer(t *testing.T, st *fakeStore, targets []hub.StorageTarget) *Server {
t.Helper()
srv, err := NewServer(Options{
ListenAddr: "127.0.0.1:0", Guests: &fakeGuests{}, Backups: &fakeBackups{}, Store: st,
Storage: fakeStorage{targets: targets},
Tokens: staticTokens{"A": 8200},
BackupTiers: []BackupTier{
{TargetID: "local", Cadence: 24 * time.Hour, Primary: true, Service: &fakeBackups{}},
{TargetID: "felhom-pbs", Cadence: 7 * 24 * time.Hour, Service: &fakeBackups{}},
},
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
})
if err != nil {
t.Fatal(err)
}
srv.baseCtx = context.Background()
srv.now = func() time.Time { return testNow }
return srv
}
// RED-PROOF (run 2026-09-15, recorded in REPORT.md): with tierBackupStates filling LastSuccess from
// pickLatestBackup(ctx, vmid, false, …) — an ATTEMPT — the pbs tier's last_success became the failed
// 0-byte record and this failed at "pbs tier reports a failed attempt as its last success".
func TestBackupStatus_TierStates_FailedTierNeverStandsInForSuccess(t *testing.T) {
st := &fakeStore{backups: []hub.Backup{
{TargetID: "local", VMID: 8200, Success: true, SizeBytes: 8877619753, StartedAt: "2026-06-10T11:03:23Z"},
{TargetID: "felhom-pbs", VMID: 8200, Success: false, Error: "storage 'felhom-pbs' does not exist", StartedAt: "2026-06-10T11:09:59Z"},
}}
srv := tierStatesServer(t, st, []hub.StorageTarget{{Name: "local", Type: "local"}}) // PBS storage ABSENT
tiers := tierStatesOf(t, srv)
if len(tiers) != 2 {
t.Fatalf("want 2 tiers, got %+v", tiers)
}
local, pbs := tiers[0], tiers[1]
if local.Target != "local" || local.LastSuccess == nil || local.LastSuccess.SizeBytes != 8877619753 || local.Storage != StoragePresencePresent {
t.Fatalf("local tier lost its successful backup: %+v", local)
}
if pbs.LastSuccess != nil {
t.Fatalf("pbs tier reports a failed attempt as its last success: %+v", pbs.LastSuccess)
}
if pbs.LastAttempt == nil || pbs.LastAttempt.Success || pbs.LastAttempt.Error == "" {
t.Fatalf("pbs tier's failed attempt is not reported as failed: %+v", pbs.LastAttempt)
}
if pbs.Storage != StoragePresenceAbsent {
t.Fatalf("pbs storage should read absent, got %q", pbs.Storage)
}
// The pre-R-517 field is unchanged (compat): still the newest record across targets.
w := do(t, srv.Handler(), "GET", "/backup/status", "A", "")
var resp struct {
Data BackupStatusResponse `json:"data"`
}
_ = json.Unmarshal(w.Body.Bytes(), &resp)
if resp.Data.Backup == nil || resp.Data.Backup.TargetID != "felhom-pbs" {
t.Fatalf("untargeted .backup changed meaning: %+v", resp.Data.Backup)
}
}
// /backup/tiers advertises storage presence, tri-state.
func TestBackupTiers_StoragePresence(t *testing.T) {
srv := tierStatesServer(t, &fakeStore{}, []hub.StorageTarget{{Name: "local"}})
w := do(t, srv.Handler(), "GET", "/backup/tiers", "A", "")
var resp struct {
Data BackupTiersResponse `json:"data"`
}
if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
t.Fatal(err)
}
got := map[string]string{}
for _, ti := range resp.Data.Tiers {
got[ti.Target] = ti.Storage
}
if got["local"] != "present" || got["felhom-pbs"] != "absent" {
t.Fatalf("storage presence wrong: %v", got)
}
// An unreadable storage view is "unknown", never "absent".
srv.storage = tierErrStorage{}
if p := srv.storagePresence(context.Background(), "felhom-pbs"); p != StoragePresenceUnknown {
t.Fatalf("unreadable storage view must be unknown, got %q", p)
}
}
type tierErrStorage struct{}
func (tierErrStorage) Observe(context.Context) ([]hub.StorageTarget, error) {
return nil, context.DeadlineExceeded
}
// A targeted request keeps the pre-R-517 bytes (no tiers array).
func TestBackupStatus_TargetedHasNoTiers(t *testing.T) {
srv := tierStatesServer(t, &fakeStore{}, []hub.StorageTarget{{Name: "local"}})
w := do(t, srv.Handler(), "GET", "/backup/status?target=local", "A", "")
if json.Valid(w.Body.Bytes()) && containsKey(w.Body.Bytes(), "tiers") {
t.Fatalf("targeted status grew a tiers array: %s", w.Body.String())
}
}
func containsKey(b []byte, key string) bool {
var m struct {
Data map[string]json.RawMessage `json:"data"`
}
_ = json.Unmarshal(b, &m)
_, ok := m.Data[key]
return ok
}
+453
View File
@@ -0,0 +1,453 @@
package localapi
import (
"context"
"encoding/json"
"os"
"path/filepath"
"sort"
"strconv"
"strings"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// R-523 — the in-guest controller supervisor.
//
// THE OUTAGE THIS EXISTS TO KILL (BIGNIGHT F9, 2026-09-14). `docker kill felhom-controller` left the
// container `Exited (137)`. Nothing restarted it: Docker never restarts a container whose stop it
// records as deliberate — measured 2026-09-15 on Docker 29.8.0 for BOTH `unless-stopped` and `always`
// (evidence-p1fixes-2026-09-15/A1) — and the golden's `felhom-controller-bootstrap.service` is a
// oneshot (`RemainAfterExit=yes`) that ran once at boot and watches nothing. The household's
// dashboard answered 502 for 33 minutes until the box was power-cycled.
//
// This is doc 03 §4's sentence made real: "Healing a crashed controller is non-destructive by
// construction … redeploy = restart … inside the existing guest — never a guest destroy." The act
// is exactly the swap's own restart (`systemctl restart felhom-controller-bootstrap.service`, which
// does `docker rm -f` + `docker run` from the baked image and the guest's persistent volume), over
// the same GuestExecutor and the same two sudoers grants (`docker inspect -f *`, the unit restart).
// No new privilege.
//
// THE GUARDS, each because doing the act at the wrong moment is worse than not doing it:
// - not during a swap (the swap stops the controller ON PURPOSE and owns its own rollback);
// - not when the operator parked it (`<guests>/<vmid>/controller-parked` on the HOST);
// - not on a guest that is not running, is locked (backup/restore/snapshot/migrate), or has a
// vzdump in flight — a stopping or restoring guest is someone else's transaction;
// - not on ONE observation: the container must be seen not-running on two consecutive sweeps, so
// the bootstrap's own rm-f/run window (boot, path-unit hot-plug) is never raced;
// - no thrash: 3 restarts inside 15 minutes → stop restarting, raise `controller_crashloop`, try
// again after 30 minutes.
//
// THE EVENTS. The agent has no event channel of its own; its heartbeat IS the channel (the
// capability/leaf precedent). The per-guest record rides the host report as `controller_supervisor`,
// and the hub's ControllerSupervisorChecker mints `controller_restarted_by_agent` (info) when a
// guest's `last_restart_at` moves and `controller_crashloop` (error, operator-only) when
// `crashloop_since` moves. Timestamps, not counters, so an agent restart (which zeroes the in-memory
// record) can never read as a new restart.
const (
// controllerSupervisorInterval is the sweep cadence. Two not-running observations are required,
// so a killed controller is restarted 30–60 s after it died.
controllerSupervisorInterval = 30 * time.Second
// controllerSupervisorConfirm is how many consecutive not-running observations license a restart.
controllerSupervisorConfirm = 2
// Backoff: controllerCrashloopMax restarts inside controllerCrashloopWindow → give up for
// controllerCrashloopPause.
controllerCrashloopMax = 3
controllerCrashloopWindow = 15 * time.Minute
controllerCrashloopPause = 30 * time.Minute
// controllerSupervisorHeartbeatEvery: a liveness line every 20 sweeps (10 minutes) — a silent
// watchdog is indistinguishable from a dead one (standing rule 3).
controllerSupervisorHeartbeatEvery = 20
// ControllerParkedMarker is the host-side file that parks a guest's controller. The operator
// creates it with `touch /var/lib/felhom-agent/guests/<vmid>/controller-parked` and removes it to
// unpark. Host-side on purpose: it needs no in-guest exec grant, it survives a guest rebuild of
// the controller container, and a customer inside the guest cannot park the supervisor.
ControllerParkedMarker = "controller-parked"
defaultGuestsStateDir = "/var/lib/felhom-agent/guests"
// R-539 (operator ruling 3 of 2026-09-16) — the SLOW crash loop. The 3-in-15-minutes brake above
// cannot see a controller that dies every 20 minutes: no two restarts share its window, so it is
// restarted for ever and the only trace is an info event that mails nobody (measured 2026-09-16,
// R-531). A second counter over 24 hours raises a WARNING at the fifth restart. It does NOT stop
// restarting — the fast brake stays the only brake, unchanged. Every restart the supervisor
// performs counts, including one that follows a deliberate operator `docker kill` (measured
// 2026-09-15: the supervisor cannot tell a kill from a crash, and a controller that is killed five
// times a day is worth a line to the operator either way).
controllerSlowCrashloopWindow = 24 * time.Hour
controllerSlowCrashloopMax = 5
// controllerSlowCounterFile holds the 24-hour restart times and the last raise, per guest, beside
// the parked marker.
controllerSlowCounterFile = "controller-restarts-24h.json"
)
// controllerSupState is one guest's supervisor record. In-memory on purpose (the guest-power
// precedent): an agent restart forgets a crash-loop pause, which costs at most one more restart
// attempt, whereas persisting it could carry a stale "give up" across the restart that fixed it.
type controllerSupState struct {
notRunningSeen int
restarts []time.Time // restart times inside the crash-loop window (pruned)
restartsTotal int
lastRestartAt time.Time
lastReason string
crashloopSince time.Time // zero = not in a crash-loop pause
parked bool
// R-539 — the slow counter. PERSISTED, unlike everything above, and the precedent's reason does not
// apply to it: persisting the fast record could carry a stale "give up" across the restart that
// fixed it, but this record never gives anything up — it only warns. Losing it on an agent restart,
// on the other hand, would hide exactly the box it exists for (one whose agent restarts too).
restarts24h []time.Time
slowCrashloopSince time.Time // the last raise; kept after it ages out, the hub keys on it MOVING
}
type controllerSupervisor struct {
mu sync.Mutex
guests map[int]*controllerSupState
sweeps int
}
// WatchControllers runs the controller supervisor sweep until ctx is done. No-op when the guest list
// (staleLock) or the guest executor is not wired.
func (s *Server) WatchControllers(ctx context.Context) {
if s.staleLock == nil || s.guestExec == nil {
s.logger.Info("controller-supervisor: not wired (no guest list or no guest executor) — disabled")
return
}
s.logger.Info("controller-supervisor: started", "interval", controllerSupervisorInterval.String(),
"confirm_sweeps", controllerSupervisorConfirm, "crashloop_max", controllerCrashloopMax,
"crashloop_window", controllerCrashloopWindow.String(),
"slow_crashloop_max", controllerSlowCrashloopMax, "slow_crashloop_window", controllerSlowCrashloopWindow.String(),
"guests_dir", s.guestsStateDir())
t := time.NewTicker(controllerSupervisorInterval)
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case <-t.C:
s.ControllerSupervisorTick(ctx)
}
}
}
func (s *Server) guestsStateDir() string {
if s.guestsDir != "" {
return s.guestsDir
}
return defaultGuestsStateDir
}
// provisionedGuest reports whether the agent provisioned a controller into this guest: the
// `<guests>/<vmid>/bootstrap` directory exists. The directory itself, not bootstrap.json inside it —
// the directory is owned by the mapped guest root (0700), so the non-root agent can see the entry but
// not stat the file within.
func (s *Server) provisionedGuest(vmid int) bool {
fi, err := os.Stat(filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), "bootstrap"))
return err == nil && fi.IsDir()
}
func (s *Server) controllerParked(vmid int) bool {
_, err := os.Stat(filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), ControllerParkedMarker))
return err == nil
}
func (s *Server) supState(vmid int) *controllerSupState {
if s.ctrlSup.guests == nil {
s.ctrlSup.guests = map[int]*controllerSupState{}
}
st := s.ctrlSup.guests[vmid]
if st == nil {
st = &controllerSupState{}
s.loadSlowCounter(vmid, st)
s.ctrlSup.guests[vmid] = st
}
return st
}
// slowCounterRecord is the on-disk shape of the R-539 counter.
type slowCounterRecord struct {
Restarts []time.Time `json:"restarts"`
SlowCrashloopSince time.Time `json:"slow_crashloop_since,omitempty"`
}
func (s *Server) slowCounterPath(vmid int) string {
return filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), controllerSlowCounterFile)
}
// loadSlowCounter restores the persisted counter into a fresh state. Absent = a clean start; unreadable
// or corrupt = a clean start with a WARN (a warning counter must never block supervision).
func (s *Server) loadSlowCounter(vmid int, st *controllerSupState) {
b, err := os.ReadFile(s.slowCounterPath(vmid))
if err != nil {
if !os.IsNotExist(err) {
s.logger.Warn("controller-supervisor: slow counter unreadable — starting it from zero", "vmid", vmid, "err", err)
}
return
}
var rec slowCounterRecord
if err := json.Unmarshal(b, &rec); err != nil {
s.logger.Warn("controller-supervisor: slow counter corrupt — starting it from zero", "vmid", vmid, "err", err)
return
}
st.restarts24h = pruneBefore(rec.Restarts, s.clock().Add(-controllerSlowCrashloopWindow))
st.slowCrashloopSince = rec.SlowCrashloopSince
if len(st.restarts24h) > 0 || !st.slowCrashloopSince.IsZero() {
s.logger.Info("controller-supervisor: slow counter restored from disk", "vmid", vmid,
"restarts_24h", len(st.restarts24h), "slow_crashloop_since", st.slowCrashloopSince.Format(time.RFC3339))
}
}
// saveSlowCounter writes the counter atomically (tmp + rename, 0600). A failure is logged and the
// in-memory counter carries on — the next restart retries the write.
func (s *Server) saveSlowCounter(vmid int, rec slowCounterRecord) {
path := s.slowCounterPath(vmid)
b, err := json.Marshal(rec)
if err == nil {
tmp := path + ".tmp"
if err = os.WriteFile(tmp, b, 0o600); err == nil {
err = os.Rename(tmp, path)
}
}
if err != nil {
s.logger.Warn("controller-supervisor: could not persist the slow counter (kept in memory)", "vmid", vmid, "path", path, "err", err)
}
}
// ControllerSupervisorTick performs one sweep. Exported so a test (and a live check) can drive one
// cycle without waiting on the ticker.
func (s *Server) ControllerSupervisorTick(ctx context.Context) {
if s.staleLock == nil || s.guestExec == nil {
return
}
guests, err := s.staleLock.Guests(ctx)
if err != nil {
// Ownership unproven ⇒ touch nothing (the guest-power rule).
s.logger.Warn("controller-supervisor: guest list unavailable — skipping sweep (ownership unproven)", "err", err)
return
}
var evaluated, down int
for _, g := range guests {
if ctx.Err() != nil {
return
}
if !s.provisionedGuest(g.VMID) {
continue
}
evaluated++
if !s.superviseOneController(ctx, g.VMID, g.Status) {
down++
}
}
s.ctrlSup.mu.Lock()
s.ctrlSup.sweeps++
sweeps := s.ctrlSup.sweeps
s.ctrlSup.mu.Unlock()
if sweeps%controllerSupervisorHeartbeatEvery == 0 {
s.logger.Info("controller-supervisor: alive", "sweeps_since_boot", sweeps,
"guests_evaluated", evaluated, "controllers_not_running", down)
}
}
// controllerRunning asks the guest's Docker for the controller's state. Returns (running, known).
// known=false means the question could not be answered (pct exec failed for a reason other than a
// missing container) — the caller does nothing on unknown. An ABSENT container is a known "not
// running": `docker rm` of the controller is the same outage as a kill.
func (s *Server) controllerRunning(ctx context.Context, vmid int) (running, known bool, status string) {
out, err := s.guestExec.GuestExec(ctx, vmid, "docker", "inspect", "-f", "{{.State.Status}}", controllerContainer)
if err != nil {
msg := strings.ToLower(err.Error() + " " + out)
if strings.Contains(msg, "no such object") || strings.Contains(msg, "no such container") {
return false, true, "absent"
}
return false, false, ""
}
status = strings.TrimSpace(out)
// "restarting" is Docker's own restart loop at work — not ours to fight on this sweep.
return status == "running" || status == "restarting", true, status
}
// superviseOneController evaluates one provisioned guest and restarts its controller when every guard
// allows. Returns false when the controller was observed not running.
func (s *Server) superviseOneController(ctx context.Context, vmid int, guestStatus string) bool {
now := s.clock()
if guestStatus != "running" {
s.resetNotRunning(vmid)
return true // the guest-power watchdog owns a stopped guest; its controller is not "down"
}
running, known, status := s.controllerRunning(ctx, vmid)
if !known {
s.logger.Debug("controller-supervisor: controller state unknown (guest exec failed) — no action", "vmid", vmid)
s.resetNotRunning(vmid)
return true
}
parked := s.controllerParked(vmid)
s.ctrlSup.mu.Lock()
st := s.supState(vmid)
st.parked = parked
if running {
st.notRunningSeen = 0
s.ctrlSup.mu.Unlock()
return true
}
st.notRunningSeen++
seen := st.notRunningSeen
s.ctrlSup.mu.Unlock()
if parked {
s.logger.Info("controller-supervisor: controller is not running and the guest is PARKED — leaving it",
"vmid", vmid, "status", status, "marker", filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), ControllerParkedMarker))
return false
}
s.swapMu.Lock()
swapping := s.swapInFlight[vmid]
s.swapMu.Unlock()
if swapping {
s.logger.Info("controller-supervisor: controller is not running during a controller SWAP — the swap owns it",
"vmid", vmid, "status", status)
s.resetNotRunning(vmid)
return false
}
if seen < controllerSupervisorConfirm {
s.logger.Info("controller-supervisor: controller observed not running — confirming on the next sweep",
"vmid", vmid, "status", status, "seen", seen, "of", controllerSupervisorConfirm)
return false
}
lock, _, err := s.staleLock.Lock(ctx, vmid)
if err != nil {
s.logger.Warn("controller-supervisor: could not read the guest lock — no action (fail-safe)", "vmid", vmid, "err", err)
return false
}
if lock != "" {
s.logger.Info("controller-supervisor: guest is LOCKED — another operation owns it, no action", "vmid", vmid, "lock", lock)
return false
}
if busy, berr := s.staleLock.BackupRunning(ctx, vmid); berr != nil || busy {
s.logger.Info("controller-supervisor: a vzdump may be in flight for the guest — no action",
"vmid", vmid, "backup_running", busy, "err", berr)
return false
}
// Backoff.
s.ctrlSup.mu.Lock()
st = s.supState(vmid)
if !st.crashloopSince.IsZero() {
if now.Sub(st.crashloopSince) < controllerCrashloopPause {
s.ctrlSup.mu.Unlock()
s.logger.Warn("controller-supervisor: crash-loop pause in force — not restarting",
"vmid", vmid, "since", st.crashloopSince.Format(time.RFC3339), "resume_after", controllerCrashloopPause.String())
return false
}
// Pause over: resume with a clean window. crashloopSince stays as the record of the last
// crash-loop (the hub keys on it moving, not on it clearing).
st.restarts = nil
st.crashloopSince = time.Time{}
}
st.restarts = pruneBefore(st.restarts, now.Add(-controllerCrashloopWindow))
if len(st.restarts) >= controllerCrashloopMax {
st.crashloopSince = now
n := len(st.restarts)
s.ctrlSup.mu.Unlock()
s.logger.Error("controller-supervisor: CRASH-LOOP — the controller would not stay up; stopping restarts and raising controller_crashloop",
"vmid", vmid, "restarts_in_window", n, "window", controllerCrashloopWindow.String(), "pause", controllerCrashloopPause.String())
return false
}
s.ctrlSup.mu.Unlock()
reason := "controller container " + status + " on " + strconv.Itoa(controllerSupervisorConfirm) + " consecutive sweeps"
s.logger.Warn("controller-supervisor: controller is NOT running — restarting the bootstrap unit",
"vmid", vmid, "status", status, "unit", bootstrapUnit)
if _, err := s.guestExec.GuestExec(ctx, vmid, "systemctl", "restart", bootstrapUnit); err != nil {
s.logger.Error("controller-supervisor: bootstrap restart failed", "vmid", vmid, "err", err)
reason += "; restart FAILED: " + err.Error()
}
s.ctrlSup.mu.Lock()
st = s.supState(vmid)
st.restarts = append(st.restarts, now)
st.restartsTotal++
st.lastRestartAt = now
st.lastReason = reason
st.notRunningSeen = 0
// R-539: the slow counter. Raise at most once per 24 hours — the hub mails on the raise MOVING.
st.restarts24h = append(pruneBefore(st.restarts24h, now.Add(-controllerSlowCrashloopWindow)), now)
n24 := len(st.restarts24h)
raised := false
if n24 >= controllerSlowCrashloopMax && (st.slowCrashloopSince.IsZero() || now.Sub(st.slowCrashloopSince) >= controllerSlowCrashloopWindow) {
st.slowCrashloopSince = now
raised = true
}
rec := slowCounterRecord{Restarts: append([]time.Time(nil), st.restarts24h...), SlowCrashloopSince: st.slowCrashloopSince}
s.ctrlSup.mu.Unlock()
s.saveSlowCounter(vmid, rec)
s.logger.Warn("controller-supervisor: RESTARTED the controller", "vmid", vmid, "reason", reason, "restarts_24h", n24)
if raised {
s.logger.Warn("controller-supervisor: SLOW CRASH-LOOP — the controller keeps dying; still restarting it, raising controller_slow_crashloop",
"vmid", vmid, "restarts_24h", n24, "window", controllerSlowCrashloopWindow.String(), "threshold", controllerSlowCrashloopMax)
}
return false
}
func (s *Server) resetNotRunning(vmid int) {
s.ctrlSup.mu.Lock()
defer s.ctrlSup.mu.Unlock()
if st := s.ctrlSup.guests[vmid]; st != nil {
st.notRunningSeen = 0
}
}
func (s *Server) clock() time.Time {
if s.now != nil {
return s.now()
}
return time.Now().UTC()
}
func pruneBefore(ts []time.Time, cutoff time.Time) []time.Time {
out := ts[:0]
for _, t := range ts {
if !t.Before(cutoff) {
out = append(out, t)
}
}
return out
}
// ControllerSupervisorStatus is the host-report stanza source (hub.ControllerSupervisorReporter).
// Nil when the supervisor is not wired, so the stanza is omitted.
func (s *Server) ControllerSupervisorStatus(_ context.Context) *hub.ControllerSupervisorStatus {
if s.staleLock == nil || s.guestExec == nil {
return nil
}
s.ctrlSup.mu.Lock()
defer s.ctrlSup.mu.Unlock()
out := &hub.ControllerSupervisorStatus{Guests: []hub.ControllerSupervisorGuest{}}
for vmid, st := range s.ctrlSup.guests {
g := hub.ControllerSupervisorGuest{
VMID: vmid,
RestartsTotal: st.restartsTotal,
LastReason: st.lastReason,
Parked: st.parked,
Crashloop: !st.crashloopSince.IsZero(),
}
if !st.lastRestartAt.IsZero() {
g.LastRestartAt = st.lastRestartAt.UTC().Format(time.RFC3339)
}
if !st.crashloopSince.IsZero() {
g.CrashloopSince = st.crashloopSince.UTC().Format(time.RFC3339)
}
now := s.clock()
g.Restarts24h = len(pruneBefore(append([]time.Time(nil), st.restarts24h...), now.Add(-controllerSlowCrashloopWindow)))
if !st.slowCrashloopSince.IsZero() {
g.SlowCrashloopSince = st.slowCrashloopSince.UTC().Format(time.RFC3339)
g.SlowCrashloop = now.Sub(st.slowCrashloopSince) < controllerSlowCrashloopWindow
}
out.Guests = append(out.Guests, g)
}
sort.Slice(out.Guests, func(i, j int) bool { return out.Guests[i].VMID < out.Guests[j].VMID })
return out
}
@@ -0,0 +1,350 @@
package localapi
import (
"context"
"encoding/json"
"errors"
"io"
"log/slog"
"os"
"path/filepath"
"strconv"
"sync"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-523 — the controller supervisor. The consequence under test is "a dead controller is started
// again", and each guard is pinned by the case where acting would be wrong.
type supExec struct {
mu sync.Mutex
status map[int]string // docker .State.Status per vmid; "" = container absent
inspectErr error // non-nil = pct exec itself failed (unknown)
restarts map[int]int
// onRestart, when set, is the status the container reaches after a restart (a crash-looper
// stays "exited").
onRestart string
}
func (f *supExec) GuestExec(_ context.Context, vmid int, args ...string) (string, error) {
f.mu.Lock()
defer f.mu.Unlock()
switch {
case len(args) >= 5 && args[0] == "docker" && args[1] == "inspect" && args[3] == "{{.State.Status}}":
if f.inspectErr != nil {
return "", f.inspectErr
}
st, ok := f.status[vmid]
if !ok || st == "" {
return "", errors.New("pct exec: exit status 1: Error: No such object: felhom-controller")
}
return st + "\n", nil
case len(args) == 3 && args[0] == "systemctl" && args[1] == "restart" && args[2] == bootstrapUnit:
if f.restarts == nil {
f.restarts = map[int]int{}
}
f.restarts[vmid]++
if f.onRestart != "" {
f.status[vmid] = f.onRestart
}
return "", nil
}
return "", errors.New("supExec: unexpected args")
}
func (f *supExec) GuestExecStdin(context.Context, int, io.Reader, ...string) (string, error) {
return "", errors.New("supExec: no stdin exec expected")
}
func (f *supExec) count(vmid int) int {
f.mu.Lock()
defer f.mu.Unlock()
return f.restarts[vmid]
}
type supClock struct{ t time.Time }
func (c *supClock) now() time.Time { return c.t }
func supServer(t *testing.T, ex *supExec, ctl *fakeGuestPowerCtl, provisioned ...int) (*Server, *supClock, string) {
t.Helper()
dir := t.TempDir()
for _, v := range provisioned {
if err := os.MkdirAll(filepath.Join(dir, strconv.Itoa(v), "bootstrap"), 0o700); err != nil {
t.Fatal(err)
}
}
clk := &supClock{t: time.Date(2026, 9, 15, 10, 0, 0, 0, time.UTC)}
s := &Server{
staleLock: ctl,
guestExec: ex,
guestsDir: dir,
swapInFlight: map[int]bool{},
logger: slog.New(slog.NewTextHandler(discardW{}, nil)),
now: clk.now,
}
return s, clk, dir
}
func runningGuest(vmid int) *fakeGuestPowerCtl {
return &fakeGuestPowerCtl{
guests: []proxmox.Guest{{VMID: vmid, Status: "running"}},
locks: map[int]string{}, onboot: map[int]bool{vmid: true},
}
}
// The consequence: a killed controller IS restarted — on the second consecutive observation, not
// the first (the bootstrap's own rm-f/run window must never be raced).
//
// RED-PROOF: delete the `systemctl restart` GuestExec call in superviseOneController → restarts
// stays 0 → "the killed controller was NOT restarted — this is R-523".
func TestControllerSupervisor_KilledControllerIsRestarted(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "running"}
s, _, _ := supServer(t, ex, runningGuest(9201), 9201)
s.ControllerSupervisorTick(context.Background())
if n := ex.count(9201); n != 0 {
t.Fatalf("restarted on the FIRST observation (restarts=%d) — must confirm on a second sweep", n)
}
s.ControllerSupervisorTick(context.Background())
if n := ex.count(9201); n != 1 {
t.Fatalf("the killed controller was NOT restarted — this is R-523 (restarts=%d)", n)
}
st := s.ControllerSupervisorStatus(context.Background())
if len(st.Guests) != 1 || st.Guests[0].RestartsTotal != 1 || st.Guests[0].LastRestartAt == "" || st.Guests[0].LastReason == "" {
t.Fatalf("report stanza did not record the restart: %+v", st.Guests)
}
// Healthy again → no further restart.
s.ControllerSupervisorTick(context.Background())
s.ControllerSupervisorTick(context.Background())
if n := ex.count(9201); n != 1 {
t.Fatalf("a running controller was restarted again (restarts=%d)", n)
}
}
func TestControllerSupervisor_AbsentContainerIsRestarted(t *testing.T) {
ex := &supExec{status: map[int]string{}, onRestart: "running"}
s, _, _ := supServer(t, ex, runningGuest(9201), 9201)
s.ControllerSupervisorTick(context.Background())
s.ControllerSupervisorTick(context.Background())
if n := ex.count(9201); n != 1 {
t.Fatalf("a removed controller container was not restarted (restarts=%d)", n)
}
}
func TestControllerSupervisor_Guards(t *testing.T) {
cases := []struct {
name string
setup func(s *Server, ex *supExec, ctl *fakeGuestPowerCtl, dir string)
}{
{"parked", func(s *Server, _ *supExec, _ *fakeGuestPowerCtl, dir string) {
if err := os.WriteFile(filepath.Join(dir, "9201", ControllerParkedMarker), nil, 0o600); err != nil {
panic(err)
}
}},
{"swap in flight", func(s *Server, _ *supExec, _ *fakeGuestPowerCtl, _ string) { s.swapInFlight[9201] = true }},
{"guest locked", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) { ctl.locks[9201] = "backup" }},
{"vzdump running", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
ctl.backupRun = map[int]bool{9201: true}
}},
{"vzdump state unknown", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
ctl.backupErr = errors.New("tasks unreadable")
}},
{"guest not running", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
ctl.guests[0].Status = "stopped"
}},
{"docker state unknown", func(_ *Server, ex *supExec, _ *fakeGuestPowerCtl, _ string) {
ex.inspectErr = errors.New("pct exec 9201: exit status 255: container not running")
}},
{"guest list unavailable", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
ctl.guestsErr = errors.New("api down")
}},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "running"}
ctl := runningGuest(9201)
s, _, dir := supServer(t, ex, ctl, 9201)
tc.setup(s, ex, ctl, dir)
for i := 0; i < 4; i++ {
s.ControllerSupervisorTick(context.Background())
}
if n := ex.count(9201); n != 0 {
t.Fatalf("guard %q did not hold: controller restarted %d time(s)", tc.name, n)
}
})
}
}
// A guest the agent did not provision (no <guests>/<vmid>/bootstrap) is never touched.
func TestControllerSupervisor_UnprovisionedGuestIgnored(t *testing.T) {
ex := &supExec{status: map[int]string{9202: "exited"}}
s, _, _ := supServer(t, ex, runningGuest(9202) /* nothing provisioned */)
for i := 0; i < 3; i++ {
s.ControllerSupervisorTick(context.Background())
}
if n := ex.count(9202); n != 0 {
t.Fatalf("an unprovisioned guest's container was restarted (%d)", n)
}
}
// No thrash: a controller that will not stay up is restarted at most controllerCrashloopMax times
// inside the window, then the supervisor raises the crash-loop and pauses; after the pause it tries
// again.
//
// RED-PROOF (run 2026-09-15, recorded in REPORT.md): with the `len(st.restarts) >=
// controllerCrashloopMax` block removed, restarts reached 10 in the first 20 sweeps and the test
// failed at "crash-looping controller restarted 10 times".
func TestControllerSupervisor_CrashloopBackoff(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "exited"}
s, clk, _ := supServer(t, ex, runningGuest(9201), 9201)
ctx := context.Background()
for i := 0; i < 20; i++ { // 10 minutes of 30 s sweeps
s.ControllerSupervisorTick(ctx)
clk.t = clk.t.Add(controllerSupervisorInterval)
}
if n := ex.count(9201); n != controllerCrashloopMax {
t.Fatalf("crash-looping controller restarted %d times in 10 minutes — want exactly %d then a pause", n, controllerCrashloopMax)
}
st := s.ControllerSupervisorStatus(ctx).Guests[0]
if !st.Crashloop || st.CrashloopSince == "" {
t.Fatalf("crash-loop not raised in the report stanza: %+v", st)
}
// Still paused 25 minutes later.
clk.t = clk.t.Add(15 * time.Minute)
s.ControllerSupervisorTick(ctx)
s.ControllerSupervisorTick(ctx)
if n := ex.count(9201); n != controllerCrashloopMax {
t.Fatalf("restarted during the crash-loop pause (restarts=%d)", n)
}
// After the pause: tries again.
clk.t = clk.t.Add(controllerCrashloopPause)
s.ControllerSupervisorTick(ctx)
s.ControllerSupervisorTick(ctx)
if n := ex.count(9201); n != controllerCrashloopMax+1 {
t.Fatalf("did not resume after the pause (restarts=%d, want %d)", n, controllerCrashloopMax+1)
}
if since := s.ControllerSupervisorStatus(ctx).Guests[0].CrashloopSince; since != st.CrashloopSince && since != "" {
t.Fatalf("crashloop_since changed without a new crash-loop: %q → %q", st.CrashloopSince, since)
}
}
// The wire shape the hub parses. The hub's controller_supervisor_test.go carries this exact JSON.
func TestControllerSupervisorStanza_WireShape(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "running"}
s, _, _ := supServer(t, ex, runningGuest(9201), 9201)
s.ControllerSupervisorTick(context.Background())
s.ControllerSupervisorTick(context.Background())
b, _ := json.Marshal(s.ControllerSupervisorStatus(context.Background()))
var m map[string][]map[string]any
if err := json.Unmarshal(b, &m); err != nil {
t.Fatal(err)
}
g := m["guests"][0]
for _, k := range []string{"vmid", "restarts_total", "last_restart_at", "last_reason", "crashloop", "parked", "restarts_24h", "slow_crashloop"} {
if _, ok := g[k]; !ok {
t.Fatalf("stanza lacks %q — the hub keys on it: %s", k, b)
}
}
}
// ---- R-539 (operator ruling 3 of 2026-09-16): the SLOW crash loop ----------------------------------
// supKillOnce kills the controller and lets the supervisor restart it (two confirming sweeps), then
// moves the clock on by gap. The container comes back "running", so each restart is a separate act.
func supKillOnce(t *testing.T, s *Server, ex *supExec, clk *supClock, gap time.Duration) {
t.Helper()
before := ex.count(9201)
ex.mu.Lock()
ex.status[9201] = "exited"
ex.mu.Unlock()
s.ControllerSupervisorTick(context.Background())
clk.t = clk.t.Add(controllerSupervisorInterval)
s.ControllerSupervisorTick(context.Background())
if ex.count(9201) != before+1 {
t.Fatalf("kill was not followed by exactly one restart (restarts %d → %d)", before, ex.count(9201))
}
clk.t = clk.t.Add(gap)
}
// The consequence: a controller that dies every 20 minutes — never three times inside the 15-minute
// brake — raises slow_crashloop on the FIFTH restart in 24 hours, and the raise does not move again on
// the sixth (the hub mails on movement; once per 24 hours is the ruling).
//
// RED-PROOF: without the slow counter the stanza never sets slow_crashloop → "five restarts 20 minutes
// apart did not raise slow_crashloop — this is R-539".
func TestControllerSupervisor_SlowCrashloop(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "running"}, onRestart: "running"}
s, clk, _ := supServer(t, ex, runningGuest(9201), 9201)
ctx := context.Background()
for i := 1; i <= 4; i++ {
supKillOnce(t, s, ex, clk, 20*time.Minute)
}
g := s.ControllerSupervisorStatus(ctx).Guests[0]
if g.Crashloop {
t.Fatalf("the 15-minute brake fired on restarts 20 minutes apart — the fixture is wrong: %+v", g)
}
if g.SlowCrashloop || g.SlowCrashloopSince != "" {
t.Fatalf("slow_crashloop raised after only 4 restarts: %+v", g)
}
supKillOnce(t, s, ex, clk, 20*time.Minute)
g = s.ControllerSupervisorStatus(ctx).Guests[0]
if !g.SlowCrashloop || g.SlowCrashloopSince == "" || g.Restarts24h != 5 {
t.Fatalf("five restarts 20 minutes apart did not raise slow_crashloop — this is R-539: %+v", g)
}
first := g.SlowCrashloopSince
supKillOnce(t, s, ex, clk, 20*time.Minute)
g = s.ControllerSupervisorStatus(ctx).Guests[0]
if g.SlowCrashloopSince != first {
t.Fatalf("the raise moved again on the 6th restart (%q → %q) — the operator would be mailed per restart", first, g.SlowCrashloopSince)
}
if !g.SlowCrashloop {
t.Fatalf("slow_crashloop cleared while the loop continues: %+v", g)
}
}
// The negative control: restarts that never reach five inside any 24 hours never raise it.
func TestControllerSupervisor_SpreadRestartsNeverSlowCrashloop(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "running"}, onRestart: "running"}
s, clk, _ := supServer(t, ex, runningGuest(9201), 9201)
for i := 0; i < 8; i++ { // eight restarts, 7 hours apart: at most 4 inside any 24 hours
supKillOnce(t, s, ex, clk, 7*time.Hour)
}
g := s.ControllerSupervisorStatus(context.Background()).Guests[0]
if g.SlowCrashloop || g.SlowCrashloopSince != "" {
t.Fatalf("restarts 7 hours apart raised slow_crashloop: %+v", g)
}
if g.Restarts24h > 4 {
t.Fatalf("restarts_24h=%d — the 24-hour window is not pruning", g.Restarts24h)
}
}
// An agent restart must not reset the slow counter (the ruling; a box whose AGENT also restarts would
// otherwise never reach five). The same state directory, a fresh Server.
//
// RED-PROOF: keep the counter in memory only → the second Server starts at 0 → "the agent restart
// reset the slow counter".
func TestControllerSupervisor_SlowCounterSurvivesAgentRestart(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "running"}, onRestart: "running"}
s, clk, dir := supServer(t, ex, runningGuest(9201), 9201)
for i := 0; i < 4; i++ {
supKillOnce(t, s, ex, clk, 20*time.Minute)
}
s2 := &Server{
staleLock: s.staleLock,
guestExec: ex,
guestsDir: dir,
swapInFlight: map[int]bool{},
logger: s.logger,
now: clk.now,
}
supKillOnce(t, s2, ex, clk, 20*time.Minute)
g := s2.ControllerSupervisorStatus(context.Background()).Guests[0]
if g.Restarts24h != 5 || !g.SlowCrashloop {
t.Fatalf("the agent restart reset the slow counter: %+v", g)
}
}
+29
View File
@@ -98,6 +98,35 @@ func (s *Server) handleRecoverOffsitePassword(w http.ResponseWriter, r *http.Req
case errors.Is(err, escrow.ErrNoEscrowBlob): case errors.Is(err, escrow.ErrNoEscrowBlob):
s.logger.Warn("local-api: offsite key recovery: the hub holds no sealed bundle for this host", "vmid", vmid) s.logger.Warn("local-api: offsite key recovery: the hub holds no sealed bundle for this host", "vmid", vmid)
writeErr(w, http.StatusNotFound, "the hub holds no sealed recovery bundle for this host — no escrow ceremony has run") writeErr(w, http.StatusNotFound, "the hub holds no sealed recovery bundle for this host — no escrow ceremony has run")
// ── R-311 (2026-08-12) — THE CODE IS RIGHT, JUST NOT FOR THE CURRENT PACKAGE. ─────────
//
// Placed ABOVE the default for the same reason ErrBundleFetch is: the default blames the
// customer, and this case is the one where the customer is provably not at fault. The code was
// used, it worked, and it opened a package the hub is deliberately keeping.
//
// 422 rather than 400: the request was well-formed AND the credential was valid — what could
// not be processed is the pairing of a correct code with the CURRENT package. A 400 would put
// it in the same bucket as a mistype, which is the whole defect. The status is the
// machine-readable half; the controller classifies on it and must never parse this sentence.
//
// The date travels in the body because it is the one fact that lets a customer recognise which
// code they are holding. No material, no code, no password — only when that package stopped
// being current, and whether it can yield a repository password at all.
case errors.Is(err, escrow.ErrCodeOpensRetained):
var ro *escrow.RetainedOpenedError
match := escrow.RetainedMatch{}
if errors.As(err, &ro) {
match = ro.Match
}
s.logger.Info("local-api: offsite key recovery: the code did NOT open the current package but DID open a RETAINED one — the customer is not at fault",
"vmid", vmid, "superseded_at", match.SupersededAt, "retained_has_restic_pw", match.HasResticPassword)
writeStatus(w, http.StatusUnprocessableEntity, false,
map[string]any{
"opens_retained": true,
"superseded_at": match.SupersededAt,
"retained_has_restic_pw": match.HasResticPassword,
},
"the recovery code is correct, but it belongs to an EARLIER sealed package (superseded "+match.SupersededAt+"), not the one currently held")
case errors.Is(err, escrow.ErrNoResticPassword): case errors.Is(err, escrow.ErrNoResticPassword):
s.logger.Warn("local-api: offsite key recovery: the bundle opened but predates the repository-password field", "vmid", vmid) s.logger.Warn("local-api: offsite key recovery: the bundle opened but predates the repository-password field", "vmid", vmid)
writeErr(w, http.StatusConflict, "the recovery code opened the bundle, but it carries NO offsite repository password (sealed before that field existed; it cannot be retro-fitted)") writeErr(w, http.StatusConflict, "the recovery code opened the bundle, but it carries NO offsite repository password (sealed before that field existed; it cannot be retro-fitted)")
+94
View File
@@ -185,6 +185,9 @@ type Options struct {
// StaleLock recovers a guest left with a stale vzdump lock by a reboot-during-backup (F2-b), run at // StaleLock recovers a guest left with a stale vzdump lock by a reboot-during-backup (F2-b), run at
// startup by RecoverStaleLockedGuests. OPTIONAL — when nil, the recovery is a no-op. // startup by RecoverStaleLockedGuests. OPTIONAL — when nil, the recovery is a no-op.
StaleLock StaleLockController StaleLock StaleLockController
// GuestsStateDir (R-523) is the agent's per-guest state dir holding <vmid>/bootstrap and the
// controller-parked marker. "" → /var/lib/felhom-agent/guests.
GuestsStateDir string
// ControllerSwapStateDir holds the per-guest swap state file (crash-safety). "" → /var/lib/felhom-agent. // ControllerSwapStateDir holds the per-guest swap state file (crash-safety). "" → /var/lib/felhom-agent.
ControllerSwapStateDir string ControllerSwapStateDir string
// Intent records drive enroll/eject intent for the self-heal watchdog (slice 10 P3). OPTIONAL — // Intent records drive enroll/eject intent for the self-heal watchdog (slice 10 P3). OPTIONAL —
@@ -362,6 +365,12 @@ type Server struct {
swapMu sync.Mutex swapMu sync.Mutex
swapInFlight map[int]bool swapInFlight map[int]bool
// R-523: the in-guest controller supervisor (controllersupervisor.go). guestExec is the same
// GuestExecutor the swap uses; guestsDir is the agent's per-guest state dir ("" → default).
guestExec GuestExecutor
guestsDir string
ctrlSup controllerSupervisor
// Network-storage verify job (SPIKE-nas-verify): the IN-MEMORY single slot + the seams the // Network-storage verify job (SPIKE-nas-verify): the IN-MEMORY single slot + the seams the
// detached pipeline runs through (tests inject; production defaults set in NewServer). // detached pipeline runs through (tests inject; production defaults set in NewServer).
netVerifyMu sync.Mutex netVerifyMu sync.Mutex
@@ -484,7 +493,9 @@ func NewServer(o Options) (*Server, error) {
s.statFile = func(path string) bool { _, err := os.Stat(path); return err == nil } s.statFile = func(path string) bool { _, err := os.Stat(path); return err == nil }
if o.ControllerSwap != nil { if o.ControllerSwap != nil {
s.swap = NewControllerSwapper(o.ControllerSwap, o.ControllerSwapStateDir, o.Logger) s.swap = NewControllerSwapper(o.ControllerSwap, o.ControllerSwapStateDir, o.Logger)
s.guestExec = o.ControllerSwap
} }
s.guestsDir = o.GuestsStateDir
return s, nil return s, nil
} }
@@ -1084,6 +1095,10 @@ type BackupTierInfo struct {
Target string `json:"target"` Target string `json:"target"`
CadenceSeconds int64 `json:"cadence_seconds"` CadenceSeconds int64 `json:"cadence_seconds"`
Primary bool `json:"primary"` Primary bool `json:"primary"`
// Storage (R-517/R-518, v0.131.0) says whether the tier's Proxmox storage exists on this host
// RIGHT NOW: "present" | "absent" | "unknown" (storage view unreadable). Additive — an older
// controller ignores it. "unknown" is never "absent": a probe failure must not skip a backup.
Storage string `json:"storage,omitempty"`
} }
func (s *Server) handleBackupTiers(w http.ResponseWriter, r *http.Request, vmid int) { func (s *Server) handleBackupTiers(w http.ResponseWriter, r *http.Request, vmid int) {
@@ -1093,11 +1108,83 @@ func (s *Server) handleBackupTiers(w http.ResponseWriter, r *http.Request, vmid
Target: t.TargetID, Target: t.TargetID,
CadenceSeconds: int64(t.Cadence.Seconds()), CadenceSeconds: int64(t.Cadence.Seconds()),
Primary: t.Primary, Primary: t.Primary,
Storage: s.storagePresence(r.Context(), t.TargetID),
}) })
} }
writeOK(w, resp) writeOK(w, resp)
} }
// storagePresence is the tri-state twin of targetStoragePresent (which must stay fail-OPEN for the
// backup path): "present", "absent", or "unknown" when the storage view cannot be read. Only a
// successful read that does not list the storage is "absent".
func (s *Server) storagePresence(ctx context.Context, target string) string {
if s.storage == nil || target == "" {
return StoragePresenceUnknown
}
targets, err := s.storage.Observe(ctx)
if err != nil {
s.logger.Warn("local-api: storage view unavailable for the tier presence report", "target", target, "err", err)
return StoragePresenceUnknown
}
for _, t := range targets {
if t.Name == target {
return StoragePresencePresent
}
}
return StoragePresenceAbsent
}
const (
StoragePresencePresent = "present"
StoragePresenceAbsent = "absent"
StoragePresenceUnknown = "unknown"
)
// TierBackupState (R-517, v0.131.0) is one tier's truth for the customer's backup page: the newest
// SUCCESSFUL backup and the last ATTEMPT, kept apart — so a failed attempt can never stand in for a
// result ("presence is not success").
type TierBackupState struct {
Target string `json:"target"`
Primary bool `json:"primary"`
Storage string `json:"storage"` // present | absent | unknown
// LastSuccess is the newest successful backup on this tier. From the in-memory record when there
// is one; otherwise from the tier's storage (after an agent restart the record is empty — the
// BIGNIGHT F2 page showed no backup at all), in which case only started_at is known and
// LastSuccessSource is "storage".
LastSuccess *hub.Backup `json:"last_success,omitempty"`
LastSuccessSource string `json:"last_success_source,omitempty"` // record | storage
LastAttempt *TierAttempt `json:"last_attempt,omitempty"`
}
// TierAttempt is the newest recorded attempt on a tier, successful or not.
type TierAttempt struct {
StartedAt string `json:"started_at"`
Success bool `json:"success"`
Error string `json:"error,omitempty"`
}
// tierBackupStates builds the per-tier view for one guest.
func (s *Server) tierBackupStates(ctx context.Context, vmid int) []TierBackupState {
out := make([]TierBackupState, 0, len(s.tiers))
for _, t := range s.tiers {
st := TierBackupState{Target: t.TargetID, Primary: t.Primary, Storage: s.storagePresence(ctx, t.TargetID)}
if b := s.pickLatestBackup(ctx, vmid, true, t.TargetID); b != nil {
st.LastSuccess, st.LastSuccessSource = b, "record"
} else if st.Storage != StoragePresenceAbsent {
if when, look := s.newestArchiveOn(ctx, t, vmid); look == archiveFound {
st.LastSuccess = &hub.Backup{TargetID: t.TargetID, VMID: vmid, Success: true,
StartedAt: when.UTC().Format(time.RFC3339)}
st.LastSuccessSource = "storage"
}
}
if a := s.pickLatestBackup(ctx, vmid, false, t.TargetID); a != nil {
st.LastAttempt = &TierAttempt{StartedAt: a.StartedAt, Success: a.Success, Error: a.Error}
}
out = append(out, st)
}
return out
}
// tierFromRequest resolves the `?target=` query parameter to a tier. // tierFromRequest resolves the `?target=` query parameter to a tier.
// //
// THE COMPATIBILITY RULE (§4): NO target parameter → the PRIMARY tier, and the echoed target is // THE COMPATIBILITY RULE (§4): NO target parameter → the PRIMARY tier, and the echoed target is
@@ -1144,6 +1231,10 @@ type BackupStatusResponse struct {
Backup *hub.Backup `json:"backup,omitempty"` // latest recorded backup for this guest Backup *hub.Backup `json:"backup,omitempty"` // latest recorded backup for this guest
// Target (R-82) echoes the tier; empty + omitted when untargeted (pre-R-82 bytes). // Target (R-82) echoes the tier; empty + omitted when untargeted (pre-R-82 bytes).
Target string `json:"target,omitempty"` Target string `json:"target,omitempty"`
// Tiers (R-517, v0.131.0) is the per-tier truth — newest success, last attempt, storage
// presence. Served on the UNTARGETED request only; additive, so an older controller reads the
// response exactly as before.
Tiers []TierBackupState `json:"tiers,omitempty"`
} }
func (s *Server) handleBackupStatus(w http.ResponseWriter, r *http.Request, vmid int) { func (s *Server) handleBackupStatus(w http.ResponseWriter, r *http.Request, vmid int) {
@@ -1155,6 +1246,9 @@ func (s *Server) handleBackupStatus(w http.ResponseWriter, r *http.Request, vmid
// across ANY target (echo == "" → pickLatestBackup's match-any path). // across ANY target (echo == "" → pickLatestBackup's match-any path).
resp := BackupStatusResponse{VMID: vmid, Phase: PhaseIdle, Target: echo, resp := BackupStatusResponse{VMID: vmid, Phase: PhaseIdle, Target: echo,
Backup: s.pickLatestBackup(r.Context(), vmid, false, echo)} Backup: s.pickLatestBackup(r.Context(), vmid, false, echo)}
if echo == "" {
resp.Tiers = s.tierBackupStates(r.Context(), vmid)
}
if job, ok := s.jobSnapshot(backupJobKey{vmid: vmid, target: tier.TargetID}); ok { if job, ok := s.jobSnapshot(backupJobKey{vmid: vmid, target: tier.TargetID}); ok {
resp.Phase = job.Phase resp.Phase = job.Phase
resp.JobID = job.JobID resp.JobID = job.JobID
+14 -2
View File
@@ -10,6 +10,8 @@ import (
"net/url" "net/url"
"strings" "strings"
"time" "time"
"gitea.dooplex.hu/admin/felhom-agent/internal/httpx"
) )
// Client is the PBS-API client for ONE PBS server. Construct with NewClient. It is pure (no // Client is the PBS-API client for ONE PBS server. Construct with NewClient. It is pure (no
@@ -31,6 +33,11 @@ type Config struct {
Secret string // token secret (from <id>.pw) Secret string // token secret (from <id>.pw)
Namespace string // PBS namespace (from storage.cfg `namespace`); "" = root. S4 per-customer tenancy. Namespace string // PBS namespace (from storage.cfg `namespace`); "" = root. S4 per-customer tenancy.
Timeout time.Duration Timeout time.Duration
// IdleConnTimeout bounds how long this client's idle keep-alive connections are retained.
// Zero means httpx.DefaultIdleConnTimeout (90s) — it does NOT mean "no timeout", which is the
// R-344 defect. Production leaves it unset; only tests set it, to avoid a 90-second wait.
IdleConnTimeout time.Duration
} }
// NewClient builds a fingerprint-pinned, token-authed PBS client. // NewClient builds a fingerprint-pinned, token-authed PBS client.
@@ -55,8 +62,13 @@ func NewClient(cfg Config) (*Client, error) {
authHeader: "PBSAPIToken=" + cfg.TokenID + ":" + cfg.Secret, authHeader: "PBSAPIToken=" + cfg.TokenID + ":" + cfg.Secret,
namespace: cfg.Namespace, namespace: cfg.Namespace,
http: &http.Client{ http: &http.Client{
Timeout: timeout, Timeout: timeout,
Transport: &http.Transport{TLSClientConfig: tlsCfg}, // R-344: this transport MUST come from httpx. pbsTargetsFromPVE (cmd/felhom-agent/
// main.go) builds a fresh Client every cycle and drops the previous one, so a
// transport with no idle timeout strands one connection per cycle, forever, on both
// sides. That leaked 388 sockets onto ep0 in 46 hours. Pinned by
// TestAbandonedClientsReleaseTheirConnections — do not inline an http.Transport here.
Transport: httpx.NewTransport(tlsCfg, cfg.IdleConnTimeout),
}, },
}, nil }, nil
} }
+173
View File
@@ -0,0 +1,173 @@
package pbs
import (
"context"
"net"
"net/http"
"net/http/httptest"
"strings"
"sync"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/httpx"
)
// connCounter is the SERVER-side observer. It counts what the server actually holds, which is the
// only thing that answers the question this file exists for: a client that believes it closed a
// connection, and a server still holding the socket, is precisely the R-344 shape. Asserting on
// anything client-side would be asserting the mechanism instead of the consequence.
type connCounter struct {
mu sync.Mutex
open int
total int // every connection ever accepted — how many times the client DIALLED
}
func (c *connCounter) hook(_ net.Conn, s http.ConnState) {
c.mu.Lock()
defer c.mu.Unlock()
switch s {
case http.StateNew:
c.open++
c.total++
case http.StateClosed, http.StateHijacked:
c.open--
}
}
func (c *connCounter) counts() (open, total int) {
c.mu.Lock()
defer c.mu.Unlock()
return c.open, c.total
}
// waitForOpen polls until the server holds want connections, or fails naming what it still holds.
func (c *connCounter) waitForOpen(t *testing.T, want int, within time.Duration, what string) {
t.Helper()
deadline := time.Now().Add(within)
for {
open, total := c.counts()
if open == want {
return
}
if time.Now().After(deadline) {
t.Fatalf("%s: after %s the server still holds %d open connection(s), want %d (%d dialled in total)",
what, within, open, want, total)
}
time.Sleep(5 * time.Millisecond)
}
}
// newCountingPBSServer is newPBSTestServer plus a ConnState hook. Kept separate rather than
// changing the shared helper, so the existing tests are untouched by this file.
func newCountingPBSServer(t *testing.T, fn http.HandlerFunc) (*httptest.Server, string, *connCounter) {
t.Helper()
cc := &connCounter{}
ts := httptest.NewUnstartedServer(fn)
ts.Config.ConnState = cc.hook
ts.StartTLS()
t.Cleanup(ts.Close)
return ts, fingerprintOf(ts), cc
}
// TestAbandonedClientsReleaseTheirConnections is Scenario A, and it is the load-bearing test for
// R-344.
//
// It models what pbsTargetsFromPVE actually does — build a client, use it once, drop it on the
// floor without closing anything — and asserts the CONSEQUENCE on the server: the connections go
// away. Before the fix every one of these stayed established forever on both sides; 388 of them
// accumulated on ep0 in 46 hours.
//
// Deliberately NOT asserted: that err == nil, or that IdleConnTimeout holds some value. Both were
// true of the leaking code.
func TestAbandonedClientsReleaseTheirConnections(t *testing.T) {
ts, fp, cc := newCountingPBSServer(t, func(w http.ResponseWriter, _ *http.Request) {
w.Write([]byte(`{"data":[]}`))
})
host, port := hostPort(t, ts.URL)
const cycles = 5
for i := 0; i < cycles; i++ {
// One fresh client per "cycle", exactly as pbsTargetsFromPVE builds one per collect.
c, err := NewClient(Config{
Server: host, Port: port, Fingerprint: fp, TokenID: "u@pbs!t", Secret: "s",
IdleConnTimeout: 50 * time.Millisecond, // production uses the 90s default
})
if err != nil {
t.Fatal(err)
}
if _, err := c.Snapshots(context.Background(), "ds"); err != nil {
t.Fatalf("cycle %d: %v", i, err)
}
_ = c // dropped here — nothing closes it, nothing can
}
if _, total := cc.counts(); total != cycles {
t.Fatalf("setup is not modelling the leak: want %d separate dials (one per abandoned client), got %d", cycles, total)
}
cc.waitForOpen(t, 0, 5*time.Second, "abandoned pbs.Clients")
}
// TestAbandonedClientsReleaseTheirConnections_ProductionDefaultIsUsable pins the value that ships.
//
// The field being settable is exactly how it could silently become zero again — and zero used to
// mean "never expire". This asserts the production path (Config leaving it unset) lands on the
// standard-library default, so the leak cannot return through an unset field.
func TestPBSClient_UnsetIdleTimeoutUsesTheDefault(t *testing.T) {
for _, tc := range []struct {
name string
cfg time.Duration
want time.Duration
}{
{"unset — the production path", 0, httpx.DefaultIdleConnTimeout},
{"explicit zero is NOT no-timeout", 0, httpx.DefaultIdleConnTimeout},
{"negative is NOT no-timeout", -time.Second, httpx.DefaultIdleConnTimeout},
{"an explicit value is honoured", 3 * time.Second, 3 * time.Second},
} {
t.Run(tc.name, func(t *testing.T) {
c, err := NewClient(Config{
Server: "pbs.example", Fingerprint: strings.Repeat("ab", 32),
TokenID: "u@pbs!t", Secret: "s", IdleConnTimeout: tc.cfg,
})
if err != nil {
t.Fatal(err)
}
tr, ok := c.http.Transport.(*http.Transport)
if !ok {
t.Fatalf("transport is %T, not *http.Transport — the httpx wiring was replaced", c.http.Transport)
}
if tr.IdleConnTimeout != tc.want {
t.Fatalf("IdleConnTimeout = %v, want %v (zero would mean connections are retained FOREVER — that is R-344)", tr.IdleConnTimeout, tc.want)
}
})
}
}
// TestPBSClient_KeepAliveStillReuses is Scenario C, and it is the guard against a "fix" that is
// worse than the bug.
//
// Disabling keep-alive entirely would also make the leak go away — by dialling a fresh connection
// for every single request, which on a box polling ~40,000 times a day is strictly worse than what
// we started with. The fix must retire IDLE connections without stopping reuse.
func TestPBSClient_KeepAliveStillReuses(t *testing.T) {
ts, fp, cc := newCountingPBSServer(t, func(w http.ResponseWriter, _ *http.Request) {
w.Write([]byte(`{"data":[]}`))
})
host, port := hostPort(t, ts.URL)
c, err := NewClient(Config{
Server: host, Port: port, Fingerprint: fp, TokenID: "u@pbs!t", Secret: "s",
IdleConnTimeout: 30 * time.Second, // long enough that reuse is what is being measured
})
if err != nil {
t.Fatal(err)
}
for i := 0; i < 3; i++ {
if _, err := c.Snapshots(context.Background(), "ds"); err != nil {
t.Fatalf("request %d: %v", i, err)
}
}
if _, total := cc.counts(); total != 1 {
t.Fatalf("one client made 3 sequential requests over %d connection(s), want 1 — keep-alive reuse is broken, which would make the poll load WORSE than the leak", total)
}
}
+30 -1
View File
@@ -283,7 +283,36 @@ func (m *Manager) Apply(ctx context.Context, fetched bool, block *hub.WirePBSDR)
h := descriptorHash(block) h := descriptorHash(block)
cf := m.loadConsumedFailed() cf := m.loadConsumedFailed()
if mk := m.loadMarker(); mk != nil && mk.Hash == h && (cf == nil || cf.Hash != h) { if mk := m.loadMarker(); mk != nil && mk.Hash == h && (cf == nil || cf.Hash != h) {
m.setStatus(&hub.PBSDRStatus{State: mk.State, StorageID: block.StorageID, Namespace: block.Namespace, AppliedAt: mk.AppliedAt}) // R-221: RE-ASSERT THE SEED, DO NOT REMEMBER IT. The marker records that this descriptor
// converged; it says nothing about whether the file the seed writes still exists.
//
// The two live in different places and die at different times. The marker is host-side
// (`<agent-state>/pbsdr/`, markerPath above); the seed's target is `agent.json`, and the
// installer's `step_agent_config` renders that file from `base = {}` unless an explicit
// `--preserve-from` is passed — it NEVER writes an `escrow` section — then replaces it with
// O_TRUNC (felhom-host-install.sh:2396, :2449, :2579; the flag is :1246, defaulting empty at
// :256). So a rebuild leaves the marker and takes the seed, the hash still matches, this
// branch returns, and `escrow.pbs_storage_id` is never written again. The customer then
// cannot run the escrow ceremony AT ALL: handleEscrowPreflight fails the `pbs_storage_id`
// row and the wizard refuses, with no way forward from inside the product.
//
// A rebuild is only the case that was measured. The same hole opens for a hand-edited or
// restored config, which is the honest reason this is a seam fix rather than an installer
// fix — the seed must be a thing the loop asserts, not a thing it did once.
//
// COST: this runs on the converged path, i.e. every tick (60 s) forever. It is one small
// file read plus a JSON parse — no exec, no network, no Proxmox call — and seedEscrowStorageID
// returns early once the value matches. That is the whole reason it is affordable here.
//
// IT MUST NEVER UN-CONVERGE THE BOX: a failure is a Warn plus a message on the published
// status, exactly as finishConverged does it. No marker write, no state change, no retry
// storm — the early return below still happens either way.
msg := ""
if err := m.seedEscrowStorageID(block.StorageID); err != nil {
msg = "escrow.pbs_storage_id seed failed: " + err.Error() + " (set it manually before the ceremony)"
m.logger.Warn("pbsdr: " + msg)
}
m.setStatus(&hub.PBSDRStatus{State: mk.State, StorageID: block.StorageID, Namespace: block.Namespace, AppliedAt: mk.AppliedAt, Message: msg})
return // idempotent: this exact descriptor already converged return // idempotent: this exact descriptor already converged
} }
+232
View File
@@ -0,0 +1,232 @@
package pbsdr
import (
"context"
"encoding/json"
"go/ast"
"go/parser"
"go/printer"
"go/token"
"io"
"os"
"path/filepath"
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// R-221 — the escrow seed must be ASSERTED on every converged tick, not remembered.
//
// These tests drive the real Apply() with a real temp-dir agent.json and a call-recording runner.
// Calling seedEscrowStorageID directly would prove nothing: the defect IS the early return in
// Apply, and a test that steps around it cannot see it.
// convergedMarker writes a marker whose hash matches the block, i.e. puts the manager on exactly
// the idempotent path where the seed used to be skipped.
func convergedMarker(t *testing.T, m *Manager, block *hub.WirePBSDR, state string) {
t.Helper()
if err := m.writeState(m.markerPath(), marker{
Hash: descriptorHash(block), State: state, AppliedAt: "2026-08-08T00:00:00Z",
}); err != nil {
t.Fatalf("write marker: %v", err)
}
}
func escrowStorageID(t *testing.T, cfgPath string) string {
t.Helper()
raw, err := os.ReadFile(cfgPath)
if err != nil {
t.Fatalf("read config: %v", err)
}
var doc struct {
Escrow struct {
PBSStorageID string `json:"pbs_storage_id"`
} `json:"escrow"`
}
if err := json.Unmarshal(raw, &doc); err != nil {
t.Fatalf("parse config: %v", err)
}
return doc.Escrow.PBSStorageID
}
// drBlock mirrors the existing valid fixture (manager_test.go descriptor()) so that validate()
// passes and Apply actually reaches the marker check — the branch these tests are about.
func drBlock(storageID string) *hub.WirePBSDR {
return &hub.WirePBSDR{
Enabled: true, StorageID: storageID, PBSTunnelIP: "10.77.0.1",
Datastore: "felhom-offsite", Namespace: "peti", TokenID: "felhom@pbs!peti",
Fingerprint: testFP,
}
}
// SCENARIO A — the seed is re-asserted on a converged box, and nothing else happens.
//
// THIS IS THE TEST THAT MATTERS. It must fail against the pre-R-221 tree; if it passes there, it is
// not testing the defect and that is the finding.
func TestSeedReasserted_OnConvergedTick_WithZeroProxmoxCalls(t *testing.T) {
r := &fakeRunner{}
st := &fakeStorage{found: true, active: []bool{true}}
c := &fakeConsumer{}
m, cfgPath := newTestManager(t, r, st, c)
block := drBlock("felhom-pbs-dr")
convergedMarker(t, m, block, "applied")
// the shape a rebuild leaves behind: escrow section present, pbs_storage_id GONE
if got := escrowStorageID(t, cfgPath); got != "" {
t.Fatalf("precondition: config already carries a storage id %q", got)
}
m.Apply(context.Background(), true, block)
if got := escrowStorageID(t, cfgPath); got != "felhom-pbs-dr" {
t.Errorf("the converged tick did not re-assert the seed: escrow.pbs_storage_id = %q, want %q.\n"+
"This is R-221: the marker survives a rebuild, the descriptor hash still matches, the early "+
"return fires and the seed never runs into the config that no longer has it — so the customer "+
"cannot run the escrow ceremony at all.", got, "felhom-pbs-dr")
}
// ...and the idempotent path is STILL idempotent. This assertion is not decorative: without it
// a "fix" that simply deletes the early return would pass the line above.
if calls := r.recorded(); len(calls) != 0 {
t.Errorf("a converged tick must execute ZERO Proxmox commands; got %d: %+v", len(calls), calls)
}
}
// SCENARIO B — an operator's own different value survives, and the warning names both.
func TestSeedReasserted_NeverClobbersAnOperatorValue(t *testing.T) {
r := &fakeRunner{}
m, cfgPath := newTestManager(t, r, &fakeStorage{found: true, active: []bool{true}}, &fakeConsumer{})
if err := os.WriteFile(cfgPath, []byte(
`{"log_level":"info","escrow":{"posture":"zero_knowledge","pbs_storage_id":"operator-chosen"},`+
`"custom_unknown":{"keep":1}}`), 0o600); err != nil {
t.Fatal(err)
}
block := drBlock("hub-chosen")
convergedMarker(t, m, block, "applied")
m.Apply(context.Background(), true, block)
if got := escrowStorageID(t, cfgPath); got != "operator-chosen" {
t.Errorf("a value a person put there was overwritten by the descriptor: got %q, want %q", got, "operator-chosen")
}
// unknown keys must still round-trip
raw, _ := os.ReadFile(cfgPath)
if !strings.Contains(string(raw), "custom_unknown") {
t.Error("an unknown config key was dropped by the re-assert")
}
}
// SCENARIO C — the ceremony preflight's live read sees the re-asserted value with NO restart.
//
// The preflight itself lives in internal/localapi and reads the file through config.Load; what this
// asserts is the half that belongs to this package: after a converged tick, THE FILE ON DISK carries
// the id, so any live re-read is green. The daemon is never restarted in this test because it is
// never started — which is the point.
func TestSeedReasserted_IsVisibleOnDiskImmediately(t *testing.T) {
r := &fakeRunner{}
m, cfgPath := newTestManager(t, r, &fakeStorage{found: true, active: []bool{true}}, &fakeConsumer{})
block := drBlock("felhom-pbs-dr")
convergedMarker(t, m, block, "applied")
// preflight's predicate BEFORE: storageID == "" → the row is NOT OK
if escrowStorageID(t, cfgPath) != "" {
t.Fatal("precondition")
}
m.Apply(context.Background(), true, block)
// preflight's predicate AFTER, from the same file the ceremony subprocess loads
if id := escrowStorageID(t, cfgPath); id == "" {
t.Error("the preflight row would still be NOT OK after a converged tick")
}
}
// A seed failure must NEVER un-converge the box: no marker rewrite, no state change, and the status
// still reports the marker's converged state — with the failure surfaced as a message.
func TestSeedReassertFailure_DoesNotUnconverge(t *testing.T) {
r := &fakeRunner{}
m, cfgPath := newTestManager(t, r, &fakeStorage{found: true, active: []bool{true}}, &fakeConsumer{})
block := drBlock("felhom-pbs-dr")
convergedMarker(t, m, block, "applied")
markerBefore, err := os.ReadFile(m.markerPath())
if err != nil {
t.Fatal(err)
}
// make the seed fail in a way it cannot recover from: unparseable config
if err := os.WriteFile(cfgPath, []byte(`{ this is not json`), 0o600); err != nil {
t.Fatal(err)
}
m.Apply(context.Background(), true, block)
after, err := os.ReadFile(m.markerPath())
if err != nil {
t.Fatalf("the marker was removed by a seed failure: %v", err)
}
if string(after) != string(markerBefore) {
t.Error("a seed failure rewrote the convergence marker — it must not touch state")
}
if calls := r.recorded(); len(calls) != 0 {
t.Errorf("a seed failure must not trigger Proxmox work; got %+v", calls)
}
st := m.Status()
if st == nil || st.State != "applied" {
t.Errorf("a seed failure must leave the box converged; status = %+v", st)
}
if st != nil && !strings.Contains(st.Message, "seed failed") {
t.Errorf("a seed failure must be surfaced on the status, got message %q", st.Message)
}
}
// SEAM WIRING — production must construct the manager with the live config path, or the whole seed
// leg is inert. Three shipped defects in this project were fully green while their seam was never
// wired, so this walks main.go's AST for the actual call rather than grepping: a commented-out call
// satisfies strings.Contains, and an AST walk cannot see a comment.
func TestProductionWiring_NewManagerGetsTheLiveConfigPath(t *testing.T) {
path := filepath.Join("..", "..", "cmd", "felhom-agent", "main.go")
fset := token.NewFileSet()
f, err := parser.ParseFile(fset, path, nil, 0) // comments not even collected
if err != nil {
t.Fatalf("parse main.go: %v", err)
}
found := false
ast.Inspect(f, func(n ast.Node) bool {
call, ok := n.(*ast.CallExpr)
if !ok {
return true
}
sel, ok := call.Fun.(*ast.SelectorExpr)
if !ok || sel.Sel.Name != "NewManager" {
return true
}
if pkg, ok := sel.X.(*ast.Ident); !ok || pkg.Name != "pbsdr" {
return true
}
// signature: (runner, px, hubc, stateDir, secretDir, configPath, logger)
if len(call.Args) < 6 {
t.Errorf("pbsdr.NewManager called with %d args, expected >= 6", len(call.Args))
return false
}
var buf strings.Builder
if err := printNode(&buf, fset, call.Args[5]); err != nil {
t.Fatalf("print arg: %v", err)
}
got := buf.String()
if !strings.Contains(got, "SourcePath") {
t.Errorf("pbsdr.NewManager's configPath argument is %q, which is not the live config path.\n"+
"With an empty or wrong path seedEscrowStorageID returns nil immediately and the entire "+
"R-221 fix is inert while every test above still passes.", got)
}
found = true
return false
})
if !found {
t.Error("no pbsdr.NewManager call found in main.go — the manager is not constructed in production")
}
}
func printNode(w io.Writer, fset *token.FileSet, n ast.Node) error {
return printer.Fprint(w, fset, n)
}
+6 -2
View File
@@ -10,6 +10,8 @@ import (
"net/url" "net/url"
"strings" "strings"
"time" "time"
"gitea.dooplex.hu/admin/felhom-agent/internal/httpx"
) )
// doer is the minimal HTTP surface the client needs; *http.Client satisfies it. // doer is the minimal HTTP surface the client needs; *http.Client satisfies it.
@@ -65,8 +67,10 @@ func NewClient(cfg Config) (*Client, error) {
timeout = 30 * time.Second timeout = 30 * time.Second
} }
hc := &http.Client{ hc := &http.Client{
Timeout: timeout, Timeout: timeout,
Transport: &http.Transport{TLSClientConfig: tlsCfg}, // R-344, consistency only: built ONCE per process, so it never accumulated and contributed
// nothing to the ep0 leak. Same missing default, corrected for the same reason.
Transport: httpx.NewTransport(tlsCfg, 0),
} }
return &Client{ return &Client{
base: strings.TrimRight(cfg.Endpoint, "/") + "/api2/json", base: strings.TrimRight(cfg.Endpoint, "/") + "/api2/json",
+9
View File
@@ -43,12 +43,21 @@ ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
SHARED_REUSE = os.path.join(os.path.dirname(ROOT), "felhom.eu", "scripts", "reuse_refs_check.py") SHARED_REUSE = os.path.join(os.path.dirname(ROOT), "felhom.eu", "scripts", "reuse_refs_check.py")
SHARED_INSTRUCTIONS = os.path.join( SHARED_INSTRUCTIONS = os.path.join(
os.path.dirname(ROOT), "felhom.eu", "scripts", "instructions_gate.py") os.path.dirname(ROOT), "felhom.eu", "scripts", "instructions_gate.py")
# R-389 — shared, like the two above: it lives in felhom.eu/scripts/ and is never copied.
SHARED_OBSERVATIONS = os.path.join(
os.path.dirname(ROOT), "felhom.eu", "scripts", "observations_gate.py")
# (label, absolute script path, args, fast) # (label, absolute script path, args, fast)
GATES = [ GATES = [
("reuse-refs", SHARED_REUSE, [ROOT], True), ("reuse-refs", SHARED_REUSE, [ROOT], True),
("instructions", SHARED_INSTRUCTIONS, [ROOT], True), ("instructions", SHARED_INSTRUCTIONS, [ROOT], True),
("published", os.path.join(ROOT, "scripts", "check-published-versions.py"), [], False), ("published", os.path.join(ROOT, "scripts", "check-published-versions.py"), [], False),
# R-273: the tag half of a release. Legs 1-2 need no network, so it runs in --fast too — the
# missing TAG is what actually broke every install, and the pre-push hook is the earliest place
# that can catch it.
("release-complete", os.path.join(ROOT, "scripts", "check-release-complete.py"), [], True),
# R-389 — a REPORT.md observation with no register row behind it. Fast: stdlib file reads.
("observations", SHARED_OBSERVATIONS, [ROOT], True),
] ]
VERDICT = {0: "OK", 1: "FAILED", 2: "INCONCLUSIVE"} VERDICT = {0: "OK", 1: "FAILED", 2: "INCONCLUSIVE"}
+41 -1
View File
@@ -95,6 +95,25 @@ PROBE_CONFIG = "configs/felhom-agent.service"
TAG_RE = re.compile(r"^v(\d+\.\d+\.\d+)$") TAG_RE = re.compile(r"^v(\d+\.\d+\.\d+)$")
# THE retention number, read from the one file that owns it. A check and the policy it enforces
# must read the same number from the same place, or they drift and the drift looks like a defect
# in something else — which is exactly what happened on 2026-08-08/09 (R-287).
RETENTION_FILE = os.path.join(os.path.dirname(os.path.abspath(__file__)), "retention-policy.json")
def retention_kept():
"""How many of the newest generic versions the registry is expected to still serve.
Fails CLOSED and LOUD: a missing or unreadable policy file makes the check INCONCLUSIVE
rather than silently unbounded. An unbounded check would re-create the red this fixed; a
silently-bounded one would be worse.
"""
with open(RETENTION_FILE, encoding="utf-8") as fh:
n = json.load(fh)["generic_versions_kept"]
if not isinstance(n, int) or n < 1:
raise ValueError("generic_versions_kept must be a positive int, got %r" % (n,))
return n
tried = [] tried = []
@@ -189,7 +208,28 @@ def main():
print(" no v<semver> tags in this repo yet — nothing to check, and nothing proven") print(" no v<semver> tags in this repo yet — nothing to check, and nothing proven")
print("\ncheck-published-versions: NOTHING TO CHECK") print("\ncheck-published-versions: NOTHING TO CHECK")
return 0 return 0
print(" %d released version(s) to verify: %s" % (len(versions), ", ".join(versions))) all_versions = versions
try:
keep = retention_kept()
except Exception as e:
inconclusive("cannot read the retention policy (%s): %s" % (RETENTION_FILE, e))
# Bound the assertion to what the registry is expected to still hold. Sorted by SEMVER, not
# lexically: "0.9.0" > "0.10.0" as strings, and that would silently drop the wrong end.
def _key(v):
return tuple(int(x) for x in v.split("."))
versions = sorted(all_versions, key=_key)[-keep:]
dropped = [v for v in all_versions if v not in versions]
print(" %d released version(s); retention policy keeps the newest %d" % (len(all_versions), keep))
print(" verifying: %s" % ", ".join(versions))
if dropped:
# NEVER silent. A bounded check that does not say what it stopped covering is how a
# narrowing becomes permanent by accident.
print(" NOT ASSERTED (older than the retention window, and therefore not expected to be")
print(" downloadable): %s" % ", ".join(dropped))
print(" ^ these versions still have git TAGS and are still installable in the sense that")
print(" their configs resolve; what is no longer asserted is the BINARY's presence.")
bad = [] bad = []
for v in versions: for v in versions:
+124
View File
@@ -0,0 +1,124 @@
#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""check-release-complete.py — the version at the head of CHANGELOG.md is a COMPLETE release.
THE DEFECT THIS IS A MACHINE FOR (2026-08-08/09, R-273). Agent v0.128.0 was built, tested,
CHANGELOG'd and published to the package registry — and its git tag was never pushed. The hub then
vouched it, and because felhom-host-install.sh fetches an agent's config files from
`raw/tag/v<version>/configs/`, EVERY fresh install and every reinstall died at step 5 of 8, as root,
on a virgin machine, for the better part of a day.
`scripts/release-agent.sh` already warns about exactly this, in as many words:
"a released version without a git tag 404s a box mid-install, as root"
The warning was there, it was correct, and the step was still missed. **So the fix is a machine and
not a reminder** — that is the whole point of this file.
WHAT IT ASSERTS, for the newest `## vX.Y.Z` in CHANGELOG.md:
1. a git tag `vX.Y.Z` EXISTS, and
2. it points at a commit that is an ANCESTOR OF (or equal to) the tip it was released from — a tag
parked on an unrelated commit is not a release, and
3. the generic package for X.Y.Z is DOWNLOADABLE.
(3) needs the network. (1) and (2) do not, and they are the half that actually failed — so this gate
is useful offline and says so rather than going quiet.
EXIT CODES, matching this repo's other gates: 0 clean, 1 convicted, 2 inconclusive. An unreachable
registry is INCONCLUSIVE for leg 3 only; legs 1 and 2 still run and can still convict.
"""
import json
import os
import re
import subprocess
import sys
import urllib.error
import urllib.request
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
GITEA_BASE = os.environ.get("GITEA_BASE", "https://gitea.dooplex.hu").rstrip("/")
OWNER, PKG = "admin", "felhom-agent"
HEAD_RE = re.compile(r"^##\s+v?(\d+\.\d+\.\d+)\b", re.M)
def git(*args):
return subprocess.run(("git",) + args, cwd=ROOT, capture_output=True, text=True)
def head_version():
ch = os.path.join(ROOT, "CHANGELOG.md")
if not os.path.exists(ch):
return None
m = HEAD_RE.search(open(ch, encoding="utf-8").read())
return m.group(1) if m else None
def main():
print("check-release-complete — the newest CHANGELOG version is a complete release")
v = head_version()
if not v:
print(" no '## vX.Y.Z' heading in CHANGELOG.md — nothing to check, and nothing proven")
return 0
tag = "v" + v
print(" newest CHANGELOG version: %s" % tag)
problems, inconclusive = [], []
# ---- leg 1 + 2: the tag, and where it points. Offline-capable. -----------------------------
r = git("rev-parse", "-q", "--verify", "refs/tags/%s^{commit}" % tag)
if r.returncode != 0:
# A shallow CI clone has no tags of its own; ask the remote before convicting, so this
# gate does not fire on a clone shape rather than on a real defect.
ls = git("ls-remote", "--tags", "origin", "refs/tags/%s" % tag)
if ls.returncode != 0:
inconclusive.append("cannot reach origin to look for tag %s: %s"
% (tag, ls.stderr.strip()[:120]))
elif not ls.stdout.strip():
problems.append(
"TAG %s DOES NOT EXIST. The installer fetches this version's configs from\n"
" %s/%s/felhom-agent/raw/tag/%s/configs/ — without the tag every install\n"
" 404s mid-run, as root. Fix: git tag -a %s <released-commit> && git push origin %s"
% (tag, GITEA_BASE, OWNER, tag, tag, tag))
else:
print(" ok tag %s exists on origin (not in this shallow clone)" % tag)
else:
sha = r.stdout.strip()
anc = git("merge-base", "--is-ancestor", sha, "HEAD")
if anc.returncode == 0:
print(" ok tag %s -> %s, an ancestor of HEAD" % (tag, sha[:10]))
else:
problems.append("tag %s points at %s, which is NOT an ancestor of HEAD — a tag parked "
"on an unrelated commit is not a release" % (tag, sha[:10]))
# ---- leg 3: the package. Needs the network. ------------------------------------------------
url = "%s/api/packages/%s/generic/%s/%s/%s" % (GITEA_BASE, OWNER, PKG, v, PKG)
req = urllib.request.Request(url, method="HEAD")
try:
with urllib.request.urlopen(req, timeout=25) as resp:
if resp.status == 200:
print(" ok package %s is downloadable" % v)
else:
problems.append("package %s returned HTTP %s at %s" % (v, resp.status, url))
except urllib.error.HTTPError as e:
if e.code == 404:
problems.append("PACKAGE %s IS NOT PUBLISHED (HTTP 404 at %s).\n"
" Fix: bash scripts/release-agent.sh %s" % (v, url, v))
else:
inconclusive.append("registry returned HTTP %s for %s" % (e.code, v))
except Exception as e:
inconclusive.append("registry unreachable (%s) — leg 3 not checked; legs 1-2 still ran" % e)
if problems:
print("\ncheck-release-complete: INCOMPLETE RELEASE")
for p in problems:
print(" - " + p)
return 1
if inconclusive:
print("\ncheck-release-complete: INCONCLUSIVE — an undetermined result is never a pass")
for i in inconclusive:
print(" - " + i)
return 2
print("\ncheck-release-complete: %s is tagged, placed and published." % tag)
return 0
if __name__ == "__main__":
sys.exit(main())
+44
View File
@@ -0,0 +1,44 @@
{
"_comment": [
"THE retention number for published agent artifacts. One file, read by everything that",
"depends on it, because a check and the policy it enforces must read the same number from the",
"same place or they drift — and the drift looks like a defect in something else.",
"",
"WHAT WENT WRONG WITHOUT IT (2026-08-08/09). The registry stopped serving felhom-agent",
"0.120.0 and older, while scripts/check-published-versions.py demanded that EVERY git tag",
"still be downloadable. Both rules are individually sensible; together they are impossible.",
"CI went red at a commit whose own run had been green the day before, on a true finding that",
"no one could act on. The red will return at the next publish unless the two read one number.",
"",
"HOW THE NUMBER WAS ARRIVED AT — stated honestly, because it is weaker than it looks.",
"generic_versions_kept is 10 because that is what the registry demonstrably holds today",
"(felhom-agent 0.121.0..0.128.0 = 10 versions, queried 2026-08-09). It is an OBSERVED state,",
"NOT a ruling anyone has been able to locate: no register row records a package prune, R-210",
"is WAITING-ON-OPERATOR and says 'Nothing was deleted; this is a list, not an action', and it",
"concerns local Docker images rather than this registry. Container packages currently hold 19",
"each, so there is no uniform ten-per-package cap visible either. See R-287.",
"",
"SO THIS FILE IS A FLOOR, NOT A LICENCE. It says: CI may assume nothing older than the newest",
"N generic versions is still downloadable. It does NOT authorise deleting anything, and the",
"operator should confirm or replace the number — at which point this file changes and both",
"readers follow it in the same commit.",
"",
"THE DEEPER BOUND, recorded so a future session does not have to re-derive it: the principled",
"limit is the hub's vouched min_agent floor (0.127.0 on 2026-08-09). Nothing can install an",
"agent below it — the hub refuses to vouch one and boxes update to the floor — so a released",
"version below the floor being un-downloadable costs nothing real. Bounding on the floor would",
"be better than bounding on a count, and it needs the gate to read the hub, which is network",
"the gate does not have today. Filed as the follow-up in R-287.",
"",
"NEVER retire a git TAG to satisfy this. felhom-host-install.sh fetches an agent's config",
"files from raw/tag/v<version>/configs/, so deleting a tag retires the ability to install that",
"version at all — a strictly worse act than an un-downloadable binary."
],
"generic_versions_kept": 10,
"readers": [
"scripts/check-published-versions.py — bounds its assertion to the newest N versions",
"documentation/runbooks/registry-retention.md (felhom.eu) — the prune procedure"
],
"recorded": "2026-08-09",
"recorded_by": "CC, from the registry's observed state; NOT from a located operator ruling"
}