Compare commits

...

34 Commits

Author SHA1 Message Date
admin 0722b2cdb0 agent: a whole-box backup that cannot fit its local target is skipped with a reason before anything starts (R-685)
gates / gates (push) Successful in 13s
Free space is read from GET /nodes/<node>/storage — GET /storage carries no usage (found live, before release).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 22:48:51 +02:00
admin 4fe2f81a32 CHANGELOG + REPORT + CONTEXT: v0.133.0 released (tag + package verified by download), not delivered
gates / gates (push) Successful in 13s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 16:25:47 +02:00
admin 9bdb4dae8f v0.133.0: a restore-test can never fill a box's disk; leftovers retried on a timer (R-672, R-673)
gates / gates (push) Successful in 13s
Space preflight before anything is created (uncompressed size from the vzdump log / PBS
snapshot, x1.2 + 5 GiB, thin metadata, off the tested guest's pool when another storage
is eligible, unknown refuses, reported as a non-pass result). Failed scratch teardown and
the stale-lock sweep retried every 10 min (the sweep under the heavy-op gate). A thin
pool crossing 90% requests an immediate host report. Six red-proofs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 16:23:15 +02:00
admin d9864a94bf CHANGELOG: v0.132.0 vouched for Day-0 installs with golden 0.246.0 (operator decision)
gates / gates (push) Successful in 16s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 12:25:49 +02:00
admin 77cd70f7c0 REPORT: agent v0.132.0 - slow crash loop, signed delivery, live proof with the production window
gates / gates (push) Successful in 13s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 11:30:45 +02:00
admin 1030abd7d6 CHANGELOG: v0.132.0 released (tag + package verified by download)
gates / gates (push) Successful in 11s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 10:27:16 +02:00
admin 18d03bd437 v0.132.0: the slow crash-loop counter (R-539, operator ruling 3 of 2026-09-16)
gates / gates (push) Successful in 12s
Beside the unchanged 3-in-15-minutes brake, a second counter: restarts the
supervisor performed in the last 24 hours. At the fifth the heartbeat stanza
sets slow_crashloop_since (moving at most once per 24 h), slow_crashloop and
restarts_24h; hub v0.117.0 mints controller_slow_crashloop (warning,
operator-only) when the timestamp moves. It never stops restarting.

Persisted per guest (tmp+rename, 0600) so an agent restart or reboot does not
reset it - unlike the fast record, whose reason for staying in memory (a
persisted give-up outliving the fix) does not apply to a counter that only
warns. Deliberate kills count. The startup line prints the new limits.

Red-proofs seen failing: no counter; the once-per-24h guard removed ('the
operator would be mailed per restart'); the save removed ('Restarts24h:1'
after an agent restart). Negative control: restarts 7 h apart never raise it.
go build/vet/test ./... green, 30 packages.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 10:26:42 +02:00
admin e98b857684 REPORT: v0.131.0 supervisor + per-tier status, delivery and live validation
gates / gates (push) Successful in 13s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 11:37:07 +02:00
admin dcdeb3d16d CHANGELOG: v0.131.0 released (tag + package verified by download)
gates / gates (push) Successful in 12s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 09:50:47 +02:00
admin 610804b98d v0.131.0: controller supervisor (R-523); per-tier backup status + tier storage presence (R-517/R-518)
gates / gates (push) Successful in 11s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 09:49:44 +02:00
admin 4586f0f7f6 re-run CI against a register that now carries R-421
gates / gates (push) Successful in 9s
The earlier run convicted correctly: instructions_gate found this repo citing R-421 while
felhom.eu's OPEN-ITEMS.md did not yet have the row. My ordering, not the gate's fault - the register
lives in felhom.eu, so a repo citing a new row must be pushed after it.
2026-09-01 12:45:52 +02:00
admin 205e22babe decoy sweep: no gate changed here, and that is the result (R-421)
gates / gates (push) Failing after 12s
All 29 gate scripts across the four repos were read and DECOYED - the label constructed without the
fact, the gate run, the verdict recorded. 16 were fooled. None of them were in this repo.

A decoy that nobody would write proves nothing, so the attempts that turned out illegitimate were
WITHDRAWN rather than counted. Both of this repo were withdrawn, and both are named in the audit.

The gates here that could not be given a plausible decoy are listed BY NAME in
felhom.eu/scripts/decoy_coverage_gate.py EXEMPT (R-426) as UNTESTED - not as sound. A gate nobody
tried to fool is UNKNOWN, and calling it sound would be the same confident guess this sweep exists
to find.

Survey table: felhom.eu/documentation/audits/AUDIT-gate-decoys-2026-09-01.md
2026-09-01 12:38:58 +02:00
admin 058b945064 gate 11: register the shared observations gate
gates / gates (push) Successful in 10s
felhom.eu/scripts/observations_gate.py, invoked across the workspace like
reuse_refs_check.py and instructions_gate.py. This repo's REPORT.md has no
observations section today, so the gate passes quietly - it is registered for
the session that writes one.
2026-08-23 13:53:20 +02:00
admin 40d857b527 CHANGELOG: the fleet runs the published bytes, not the proof build (R-349)
gates / gates (push) Successful in 7s
Both boxes were first given a hand build: same source, same version
string, different bytes (256e0829 vs the published a56a92a7), because
release-agent.sh builds with -trimpath -buildvcs=false and a hand build
does not.

Nothing would have corrected it. The boxes already reported 0.130.0, so
self-update saw the vouched version as installed and would have done
nothing, forever. Every version check in the system compares the STRING.

Both reinstalled from the downloaded package; both now report a56a92a7.
2026-08-20 12:53:37 +02:00
admin 7ae6990bac v0.130.0 released: tag + package published, heading now claims it
gates / gates (push) Successful in 8s
sha256 a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3
tag v0.130.0 at 7569f34, 14,141,158 bytes.

Reproducible: rebuilding with -trimpath -buildvcs=false matches the
published artifact byte for byte (R-186's property, checked not assumed).

The heading said UNRELEASED while the fix was hand-installed on demo-hp
only -- publishing then would have pushed it onto the control box through
self-update. release-complete convicted on the release heading and was
right to; the answer was to stop claiming a release, not to bypass it.

NOT VOUCHED by this commit. Vouching is the separate operator act.
2026-08-20 12:47:25 +02:00
admin 7569f34aeb CHANGELOG: correct a claim this session made and then disproved
gates / gates (push) Successful in 8s
The v0.130.0 draft said the 388 descriptors already stuck on ep0 would
persist until the PBS proxy restarted. Measured within the hour: they
clear when the AGENT restarts. demo-hp released its 199 in one second
(415 -> 216 fd); demo-felhom released the remaining 203 (220 -> 17 fd in
under two seconds). 17 is precisely ep0's t0 baseline of 2026-08-18.

They were held on both sides. Closing either side ends them. ep0 was
read-only throughout and its proxy PID never changed.
2026-08-20 12:39:12 +02:00
admin ede49b610d R-344: restore the idle-connection timeout our hand-rolled transports lost
gates / gates (push) Successful in 7s
Every client here pins TLS, so none can use http.DefaultTransport and each
hand-rolls its own. A composite literal takes IdleConnTimeout ZERO, which
means retain idle connections forever -- not "use a sane default".
pbsTargetsFromPVE builds a fresh pbs.Client every cycle and drops the
previous one, and an abandoned http.Transport does not close its
connections. One stranded socket per cycle, on both sides, forever.

Measured: 388 established connections on ep0 over 46 h, 194 per box, zero
closed in a 31-minute window. pvestatd and proxmox-backup-client made
162,404 requests in the same window and leaked none.

New leaf package internal/httpx owns the default (90s, http.DefaultTransport's
own value) and NewTransport, which returns a FRESH transport per call and
treats <=0 as "use the default", never "no timeout". pbs.Config gains
IdleConnTimeout for tests only.

hub and proxmox carried the same missing default and are corrected here for
consistency. Neither contributed to the ep0 leak -- both are built once per
process and neither talks to ep0:8007.

Tests count connections SERVER-side and model the abandonment, so they pin
the consequence rather than the field. Two red-proofs, both seen failing:
removing the timeout -> "still holds 5 open connection(s), want 0";
DisableKeepAlives -> "3 sequential requests over 3 connection(s), want 1"
(the leak test PASSES under that one -- it is the worse-fix guard that
catches it).

Not released: hand-installed on demo-hp only so demo-felhom stays the
control. CHANGELOG heading stays UNRELEASED until the publish is authorised.
2026-08-20 11:09:39 +02:00
admin f17ed11599 REPORT: agent v0.129.0 — the retained-package recovery class, released and deployed
gates / gates (push) Successful in 21s
2026-08-12 18:49:51 +02:00
admin 1db56bf837 v0.129.0 — a correct code for an earlier package stops being called wrong (R-311)
gates / gates (push) Successful in 14s
Yesterday's drill proved a retained escrow package opens a set-aside store and
restores planted files byte-identical, while this agent answered the customer's
correct code with "the recovery code did not open the sealed bundle". Nothing had
ever tried the retained packages, so a correct-but-earlier code and a mistype were
genuinely indistinguishable.

OffsiteKeyRecoverer gains an optional FetchRetained, consulted ONLY after the
current package refuses, so the ordinary recovery pays nothing for it and cannot
fail because of it. A match returns ErrCodeOpensRetained wrapped in a
RetainedOpenedError carrying the supersession date - no material, no code, no
password. The local API answers 422: a FIFTH status added to the R-224 switch,
never a restructuring of it.

Fail-safe in every direction. Nil fetcher, a hub too old for the route (404 is a
clean "none"), a transport failure, a malformed package: each leaves the original
refusal standing. Attempts bounded at 6 because each unwrap is ~1s of scrypt.

Seven tests with REAL age crypto - the two situations are indistinguishable AT
THE UNWRAP, so a faked unwrap would prove nothing. Red-proof asserted applied:
remove the retained lookup and the fail-closed wrong-code error returns, which is
the lie in those exact words.
2026-08-12 18:40:00 +02:00
admin 53d047a6c1 Two guards, one number: bound the published check to the retention it must live with
gates / gates (push) Successful in 17s
Gates only. No release, no version bump, no binary published; the agent stays
v0.128.0 at 28ba8593b8 and nothing on a customer's machine changes.

THE COUPLING DEFECT. The registry stopped serving 0.120.0 and older while
check-published-versions.py demanded every tag still be downloadable. Both rules
are sensible and together they are impossible, so CI went red at a commit whose
own run had been GREEN the day before -- and would have gone red again at the
next publish when 0.121.0 was evicted. scripts/retention-policy.json is now THE
number and both readers take it from there.

WHAT CI NO LONGER COVERS, and it prints this on every run rather than leaving it
to be discovered: a released version older than the retention window is no longer
asserted downloadable. Its git tag and its config tree ARE still asserted -- only
the binary's presence is dropped. A missing policy file is INCONCLUSIVE (exit 2),
never silently unbounded.

THE NUMBER IS NOT A LOCATED RULING and the file says so in its own header. Ten is
what the registry demonstrably holds; no register row records a prune, R-210 says
"Nothing was deleted; this is a list, not an action" and concerns local Docker
images, and container packages hold 19 each. The principled bound is the hub's
vouched min_agent floor -- nothing can install below it -- and that is the
recorded follow-up.

check-release-complete.py is the tag half as a machine. release-agent.sh already
warned that "a released version without a git tag 404s a box mid-install, as
root" and the step was still missed, so this is a gate and not a reminder. Legs
1-2 need no network and run in --fast, so the pre-push hook is the earliest
catch. Red-proved by repointing the CHANGELOG head at an unreleased v0.129.0:
both legs convicted and each named its fix command.

Three controls run: green at 10 naming what it dropped; widened to 11 the evicted
version re-enters and convicts; policy removed gives INCONCLUSIVE naming the path.
2026-08-09 19:05:05 +02:00
admin 28ba8593b8 v0.128.0 — R-221: the escrow seed is asserted every tick, not remembered once
gates / gates (push) Failing after 28s
A rebuilt box could not run the escrow ceremony AT ALL, with no way forward from inside the product.
This was the only open item blocking a customer from something we promise them.

MECHANISM, established at file:line rather than assumed. The preflight refuses on
escrow.pbs_storage_id; the pbsdr bridge writes that key; and it wrote it from exactly one place —
finishConverged, reached only on the paths that actually converge.

The marker and the key live in different places and die at different times. The marker is host-side
(<agent-state>/pbsdr/marker.json). The key is in agent.json, which step_agent_config renders from
`base = {}` unless an explicit --preserve-from is given (felhom.eu/scripts/felhom-host-install.sh:
2396 the step, :2449 the render, :2579 the O_TRUNC write; the flag :1246, defaulting empty at :256)
— AND THE RENDER NEVER WRITES AN escrow SECTION AT ALL (grep over the whole heredoc: zero hits). So
a rebuild keeps the marker and takes the key: same descriptor, same hash, early return, and the seed
never runs again into a config that no longer has it.

A rebuild is only the case that was measured. The same hole opens for a hand-edited or restored
config, which is the honest reason this is a seam fix rather than an installer fix: the seed must be
a thing the loop ASSERTS, not a thing it did once.

Apply now re-asserts the seed BEFORE the idempotent early return. seedEscrowStorageID is unchanged
and still never clobbers a different existing value — an operator's own choice outranks the
descriptor's, with a warning naming both.

THE EARLY RETURN IS KEPT. It stops a converged box re-running Proxmox operations every 60s, and
TestSeedReasserted_OnConvergedTick_WithZeroProxmoxCalls asserts ZERO recorded runner calls on that
tick, so a "fix" that simply deleted the return would fail. Cost: one small file read plus a JSON
parse per tick, no exec, no network, early-returning once the value matches.

A seed failure can never un-converge the box: Warn plus a message on the published status, exactly
as finishConverged does it — no marker write, no state change.

Tests drive the REAL Apply with a real temp-dir agent.json and a call-recording runner; calling
seedEscrowStorageID directly cannot see the early return, which IS the defect. Production wiring
(pbsdr.NewManager(..., cfg.SourcePath, ...)) is asserted by walking main.go's AST, not by
strings.Contains, which a commented-out call also satisfies.

Red-proofs, each with the mutation asserted applied: removing the new call makes Scenario A fail on
today's tree (it did, with the intended message); removing the early return makes the
zero-Proxmox-calls assertion fail (it did).

go build / go vet / go test ./... green (29 packages), run separately from this commit.
2026-08-08 16:29:13 +02:00
admin 6981450110 docs: a comment claimed the hub reads a field it has no field for (R-260)
gates / gates (push) Successful in 26s
Comment-only; no behaviour, no wire change, no version bump, nothing to rebuild.

HostReport.SelfUpdatePending / SelfUpdatePendingVersion carried "The hub reads an absent field as
pending=false, the correct default." The hub has NO FIELD for either, so it reads nothing — present
or absent — and encoding/json discards them on arrival. The sentence described an intent rather than
the code and read as settled long enough that a class sweep had to find it.

The emission is correct and stays: the agent reports the truth and the fault is entirely in the
receiving. The missing consumer is R-264 (OPEN). felhom.eu/scripts/wire_contract_gate.py now refuses
any NEW field of this shape and records the existing ones as reasoned allowlist entries.
2026-08-08 08:46:17 +02:00
admin 703db166e7 v0.127.0: a mount Felhom itself made is not 'something else' (R-220)
gates / gates (push) Successful in 8s
After a rebuild the customer's own drives could not be re-attached: candidates
returned initialize:[] attach:[] while both drives sat there, and the deploy
refused with 'choose an attached drive from the list' — a list that was empty.
Measured live three times.

Mechanism: enrolment mounts a drive TWICE, at /mnt/felhom-drives/<name> and at
the raw /mnt/<name> it creates on the host. The host survives a guest rebuild;
the controller's registry does not. So classifyClaim saw a mount outside the
managed prefix and concluded 'claimed by something else' — about our own mount.

The fix is CORROBORATED, not a widened prefix: a non-managed mountpoint is
forgiven only when the SAME device is also mounted under the managed path, a
pairing only our enrolment produces. A disk another system uses — /srv/data,
/media/x, even /mnt/someone-elses-disk — has no counterpart and is STILL
refused, with its own test and a red-proof showing an over-wide fix offering it
for formatting.

Read from /proc/mounts deliberately: the lsblk invocation is pinned verbatim in
configs/felhom-agent.sudoers, so using the plural MOUNTPOINTS would have coupled
this to a sudoers rollout. /proc/mounts is world-readable — no sudo, no new
allowlisted command, no config change.

Fail-safe: an unreadable mount table corroborates NOTHING, so the device
classifies exactly as before. 'Could not corroborate' must never read as 'ours'.

29 packages ok, vet clean, agent gates OK.
2026-08-06 12:55:28 +02:00
admin aa74294a7d docs: felhom-agent CLAUDE.md becomes a core plus path-scoped rules (R-229 leg b)
gates / gates (push) Successful in 8s
175 -> 99 effective lines. New .claude/rules/{proxmox,localapi,backup,storage}.md alongside the
existing health-checks.md. The release section points at the felhom-build-deploy skill rather than
restating a table that drifts from the script; the layout section's per-package annotations moved
into the rule file for their area instead of being deleted.

Kept in the core because it is the only part re-injected after /compact: the root-CLI fence and its
three exceptions, the destructive-op gate, prove-ownership (audit A1), the gate entry point, the F9
live-validation fence, and the checklist.

health-checks.md overlaps localapi.md and storage.md on three globs -- deliberate, both load,
stated in each file. Go build/vet/test green and unchanged.
2026-08-06 11:28:27 +02:00
admin 5b2666e3a2 docs: R-168 is CLOSED — the "CI is still owed" sentence was stale (R-229 part 2)
gates / gates (push) Successful in 11s
Corrected in all four instruction files across all four repos. Found while confirming this
session own push by run ID, which is precisely the check that catches it.

In felhom-agent/CLAUDE.md the sentence contradicted the same file release section, which
already said R-168 mails the failure -- a contradiction inside one instruction file, the exact
class the R-229 work exists to find.

REPORT.md deliberately NOT overwritten in the sibling repos: a one-line docs correction must not
destroy the record of their last real implementation.
2026-08-06 11:02:59 +02:00
admin 062a7027ab docs: remove expired and contradictory blocks from CLAUDE.md (R-229)
gates / gates (push) Successful in 8s
Surgical corrections only; the file is deliberately NOT restructured (deferred).

Deleted the expired TEMPORARY block. It read "felhom-pve is at a remote site
(until ~2026-08-02) ... Delete this block on return" and was still being read as
current fact on 2026-08-06, four days past its own deadline, while
felhom-controller/CLAUDE.md asserted the opposite. The location-independence fact
worth keeping (localapi binds 169.254.253.1:8443 on vmbr9 since the R-50 island
migration) moved to an HTML comment.

Every component version literal is gone from effective text, including the
--version reading and the go.mod Go directive. Versions change several times a
day; ask the hub's /hosts + /configs or the box.

The drill-VM claim and the host addresses now point at
documentation/operations/nodes.md, which already stated both correctly. This
file's drill-VM claim was the correct one -- confirmed by qm list on demo-hp.

The R-115/R-188/R-186 release narratives moved to an HTML comment and to the
felhom-build-deploy skill; the directives stayed (never hand-roll the build; the
build -> tag -> publish -> push order; reproducible -trimpath -buildvcs=false).

The health-check block-I/O rule became .claude/rules/health-checks.md, scoped to
the five packages where health checks are written. It had been duplicated from
felhom.eu/CLAUDE.md with a note explaining that that file does not load in an
agent-only session -- correct reasoning, made obsolete by path-scoped rules.

agent_gates.py registers the shared instructions gate.

Docs only -- no Go, no version bump, nothing built or deployed.
Ledger: felhom.eu/documentation/audits/LEDGER-instruction-trim-2026-08-06.md

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JJc8sAGRWmavP3rMtdpkr2
2026-08-06 09:38:38 +02:00
admin a2e914f683 v0.126.0: a fetch failure is not a wrong recovery code (R-224)
gates / gates (push) Successful in 7s
A hub the agent could not reach was reported to the customer as a bad recovery
code. Measured live 2026-08-05 (CAMPAIGN-11 F3): hub firewalled off, a CORRECT
current code, and the customer told it did not open their package — in 0.0556s
against ~1.0s for a real unseal. No unseal was attempted.

The discriminator existed here and this boundary threw it away: recover.go
fails at four distinguishable points and the local-api handler had cases for
two, with a default answering 'the recovery code did not open the sealed
bundle, OR the bundle could not be fetched'.

escrow.ErrBundleFetch now joins the fetch leg and the handler routes it to 502
with its own words — the code was NOT used. 502 not 4xx: the request was not
bad, an upstream dependency failed. Four situations, four statuses: 502 fetch /
400 fetched-and-refused / 404 no bundle / 409 predates the field. The
controller classifies on the STATUS and never parses the sentence.

A GREEN TEST NAMED THIS DEFECT AND DID NOT PREVENT IT.
TestRecoverOffsiteRepoPassword_FetchErrorIsDistinct has said since v0.125.0
that the operator must not be sent to re-read their code because the hub was
unreachable — and passed throughout, because it asserted this package's error
STRING one layer below the merge, and a string is not something a caller can
branch on. Re-pointed at the sentinel, with a consequence-level twin asserting
the status.

Red-proofs: removing the %w join fails the sentinel test; deleting the handler
case makes fetch and wrong-code both answer 400 with the wrong-code sentence.

29 packages ok, vet clean, agent gates OK.
2026-08-06 07:55:15 +02:00
admin 0404f60e6a pre-push: refuse a push from a clone outside the felhom workspace (R-204 rider)
gates / gates (push) Successful in 10s
The workspace root is already documented (workspace-CLAUDE.md, the workspace-root
CLAUDE.md 'stay inside it') and work drifted into a home directory anyway. A rule
that has failed once as a reminder is not fixed by writing it down again, so it is
now asserted where it can bite.

A push is the right trigger: throwaway clones under /tmp for probes and red-proofs
never push, so nothing legitimate breaks. Symlinks are resolved on both sides; an
absent workspace root SKIPS the check rather than failing it, so this cannot brick
a legitimate clone on another machine. The only bypass is the documented
--no-verify, whose use is already reportable.

Identical in all four repos.
2026-08-05 10:46:39 +02:00
admin 3f5f61b716 docs: R-199 links 6-8 — CONTEXT + REPORT (proven live on demo-felhom)
gates / gates (push) Successful in 7s
2026-08-04 13:56:38 +02:00
admin 6d7904786c agent v0.125.0: open the sealed bundle, return one field (R-199 links 7-8)
gates / gates (push) Successful in 7s
Link 7's only production caller was a --selftest reading R from an env var. Link 8 did not
exist: that selftest writes the whole bundle JSON and its success message named
"tunnel_token + pbs_token" -- accurate when written, a misstatement since v0.77.0 sealed the
offsite repository password into the same bundle. It now names what THIS bundle carried and
what it did not.

POST /escrow/recover-offsite-password: the controller supplies R, the agent fetches this
host's own blob from the hub (self-scoped by the per-host key), unseals it, and returns ONLY
the offsite restic repository password plus its sha256. Not the tunnel token, not the PBS
token, not the WG key -- the controller is a trust tier down and needs none of them.

R: in memory for one call, cleared on every path, never on disk, never in argv, never logged,
never echoed. A test redirects TMPDIR and asserts the tree is EMPTY afterwards -- emptiness
rather than a content scan, because a content scan is defeated by a later call overwriting the
leaked file, which is how the first version of that test passed its own red-proof while R sat
on disk.

Three distinct outcomes: no blob (404), a bundle that opens but predates the field (409), a
code that does not open it (400, fail-closed at the KDF, nothing written).

The wiring is asserted by an AST walk from func main() to the Options field, not by grep.
2026-08-04 13:41:12 +02:00
admin 856a127cd6 v0.124.1: the repair record must survive the probe that did NOT feed the hub (R-190)
gates / gates (push) Successful in 6s
v0.124.0's transition record never reached the hub, and only the live run showed
it. The capability reported degraded for "one cycle" — the probe call that did the
repair. But probeAll is invoked independently by the self-check log and by the
collector building a host report. On the demo box the repairing call was the log's
(09:39:34, journal shows the self-repair and degraded=1) and the report three
seconds later found the grant present and sent ok. The agent's journal had the
record; the hub had nothing. That is the silence R-190 is about, re-created inside
its own mitigation, with every unit test green.

Fixed with a latch on TIME, not call count: a confirmed repair reports for 20
minutes, which exceeds the 900s report interval, so at least one report must carry
it. It clears on its own and is per tier.

Two hollow tests caught and fixed on the way — one asserting a value it built
itself, one asserting the latch helper rather than the path consuming it (its
red-proof duly passed). The decisions now live in storeGrantHealthyVerdict and
storeGrantRepairedVerdict and the tests call those.
2026-08-04 09:44:56 +02:00
admin 257c4d85c0 v0.124.0: a lost storage grant repairs itself, and says that it was lost (R-190)
gates / gates (push) Successful in 7s
R-190 is a grant that worked at 04:44 on 2026-08-03 and was gone by 09:24, with a
reinstall, logged pveum activity and cluster-log entries all ruled out. The cause
is open; the resilience need not wait for it.

Everything needed already existed and had only ever been called once: the root
wrapper's `grant` verb, its sudoers vector (`grant *`, any storage id — confirmed,
not assumed), and the exact command. The verb had only ever run at storage
creation — the "built but never wired" shape in a verb rather than a seam.

The probe now runs that wrapper on a missing grant and re-reads ONCE to confirm,
the pbsdr R-22 shape including its restraint.

The record is the half that matters. A repair leaving only "ok" behind destroys
the only evidence a permission vanished, so a recurring loss becomes undetectable
— worse than the fault. A confirmed repair therefore reports DEGRADED for exactly
one cycle with the explanation in Feature, because that is the field the hub puts
in the operator's email (Reason does not travel). Nothing new was built: the hub's
existing ok->degraded->ok edge is the channel, so one loss produces one alert pair.
No wire change, no hub change, no new event type.

Bounded at one attempt per tier per hour: a storage can be unreadable for reasons
an ACL cannot fix, and re-granting every cycle is a repair loop wearing a fix's
clothes. A failed repair never masks the fault.
2026-08-04 09:38:27 +02:00
admin 72161f6cf0 REPORT: correct the manifest commit hash (311dc06)
gates / gates (push) Successful in 7s
2026-08-03 19:04:44 +02:00
admin 03b58cec0a REPORT + CONTEXT + REUSE: R-185 closed, with the corrected root cause
gates / gates (push) Successful in 7s
The installer defect was NOT PVE_STORAGES as the row and the task assumed: the
create arm of configure_backup_target grants, the Scenario-F reuse arm did not.
Also records the measured trap (an ungranted path answers with INHERITED
privileges, not empty and not 403), the deviation from the spec's suggested
Prober generalisation in favour of the existing poolReadStatus precedent, the
hollow test caught before it shipped, and that demo-hp carried the same drift and
was fixed.
2026-08-03 19:04:15 +02:00
62 changed files with 6326 additions and 524 deletions
+46
View File
@@ -0,0 +1,46 @@
---
paths: ["internal/backup/**", "internal/pbs/**", "internal/pbsdr/**", "internal/dr/**"]
---
# Backup, PBS and DR
`internal/backup/` is the vzdump runner, restore-test scheduler and report store. `internal/pbs/` is
the fingerprint-pinned PBS-API client plus the verify maintenance loop. `internal/pbsdr/` and
`internal/dr/` carry the DR tier and recipe halves.
## The three PBS laws
1. **Set-only.** `pvesm remove` **DELETES the encryption key**. Re-apply configuration; never remove
and re-add a PBS storage to change it.
2. **Secret on stdin.** A token secret is passed on stdin, never as an argv the process table shows.
3. **Verify the pin BEFORE consuming the secret.** A fingerprint check after the secret has been sent
protects nothing.
## Verify is server-side, and its default skips the work
The agent drives verification **remotely** via the PBS API; `proxmox-backup-client` has **no** verify
subcommand. `POST .../verify` defaults to **`ignore-verified=true`, which SKIPS already-verified
snapshots** — send `ignore-verified=false` to actually re-read and detect corruption. A verify that
skipped everything reports success.
## Presence is not success
A timestamp recording an **attempt** is not evidence of a **result**. Where a status field travels
beside a timestamp, the verdict must consult **both** — or the timestamp must record only successes.
Ask of any timestamp: *what exactly must have happened for this to be set?* If the answer is "we
tried", it cannot answer "did it work".
**Corollary:** when a verdict changes which field it counts from, the alarm text changes with it.
Leaving a message reading `last run 8h ago` while alarming on a six-day-old **success** turns a true
alarm into one the operator dismisses.
## Prune is server-side now
`DatastoreBackup` carries **no** `Datastore.Prune`. Boxes set `keep_last: 0` and the off-site endpoint
runs the prune jobs. **Box tokens stay write-only — never widen that grant** (R-89).
<!--
The ignore-verified default is the sharpest instance of the "absent log line" class in this repo: a
verify that silently skipped every snapshot completes fast, exits clean, and reports the same shape
as one that read every byte.
-->
+26
View File
@@ -0,0 +1,26 @@
---
paths: ["internal/capability/**", "internal/storage/**", "internal/localapi/**", "internal/hub/**", "internal/guesthook/**"]
---
# A health check issues no block I/O
No `statfs`, no `getdents`, no read, write or `fsync` — **not even behind a timeout**.
A probe that touches a wedged device enters uninterruptible sleep, survives `SIGKILL`, and cannot be
recovered until the device returns or the host reboots — so `systemctl restart` hangs too. A timeout
protects the caller's control flow and nothing else: the blocked thread remains.
**Liveness is decided from `/proc` and the kernel's own state**, never by reading or writing the
filesystem.
<!--
Measured, R-117 spike §6.3 (felhom.eu/documentation/audits/SPIKE-r117-bind-liveness-2026-07-30.md):
a probe stayed in D state 3m50s after kill -9; a buffered write with no fsync blocked too (O_CREAT
needs journal access); and statfs/getdents returned HEALTHY on a namespace that EIOs every byte —
fast, and wrong.
This rule used to be duplicated verbatim in felhom-agent/CLAUDE.md with a note explaining that
felhom.eu/CLAUDE.md "does not load in an agent-only session". That reasoning was correct before
path-scoped rules existed. The single source is now felhom.eu/CLAUDE.md "Code quality rules"; this
file is the scoped copy that loads exactly where health checks are written. (2026-08-06)
-->
+44
View File
@@ -0,0 +1,44 @@
---
paths: ["internal/localapi/**", "internal/authz/**", "internal/guesthook/**"]
---
# Local API, authz and guest hooks — the per-guest blast radius
`internal/localapi/` is the narrow per-guest local API: token store, disks/format, guest binds,
controller swap, stale-lock recovery, pinned self-signed leaf. `internal/authz/` is the operator
signed-op verifier (SSHSIG) plus the durable nonce store. `internal/guesthook/` installs the
pre-start self-heal hookscript.
> **Overlap note:** `health-checks.md` also matches `internal/localapi/**` and
> `internal/guesthook/**`. That is deliberate — both rules apply there and both load. Neither
> supersedes the other.
## Scoping is the whole security property
This API is reachable **from inside a customer guest**. Every route must be scoped to the guest that
called it — a route that can name another guest's id has escaped its blast radius. Fail **safe to
protected**: an unrecognised or unresolvable caller gets less access, never more.
## Replay protection must survive a restart
**`authz.MemoryNonceStore` on a real host is a defect** — replay protection dies on restart. Use
`authz.FileNonceStore`. The memory store exists for tests.
## The token is a hash on disk, plaintext only at mint
The store keeps **hashes**. The plaintext token exists in exactly one place, `bootstrap.json` on the
PVE host — so a "read the token" step means reading that file, and a lost token is re-minted, never
recovered.
## Binds can brick guest boot
| Do not | Because | Use |
|---|---|---|
| `GuestBinder.AttachBind`/`DetachBind` (per-drive `pct set -mpN`) | legacy model; a missing bind source can **brick guest boot** (C1) | `AttachDrive`/`DetachDrive` (intermediary model) |
| `isHostMountpoint` to reconcile bind state | a boolean cannot converge stacked double-binds (the `/mnt` doubling bug) | `countHostMounts` normalization inside `AttachDrive` |
<!--
Why fail-safe-to-protected rather than fail-closed: this API also carries the recovery paths. A hard
refusal on an unresolvable caller would make a half-broken guest unrecoverable through the very
interface built to recover it. Less access, never none.
-->
+44
View File
@@ -0,0 +1,44 @@
---
paths: ["internal/proxmox/**", "internal/reconcile/**", "internal/signedjobs/**"]
---
# Proxmox — the API contract, and how destructive work is gated
`internal/proxmox/` is the API-first `Client` plus the fenced root-CLI `Privileged`.
`internal/reconcile/` is the reconcile engine, reversibility gate, op journal and crash recovery.
`internal/signedjobs/` holds the operator-signed destructive executors (wipe, decommission).
## A 200 on the POST is not success
**Every mutating op is async**: it returns a **UPID**, and `WaitTask` must assert
`exitstatus == "OK"`. Authorization can fail at *task execution* long after the HTTP call returned
200. Treating the POST's status as the result is how a failed destroy reads as a successful one.
## The privsep token gotcha
A `--privsep 1` token's rights are the **intersection** of the backing user's permissions **and** the
token's own ACLs. The role must be granted on **both** or every call 403s. The same intersection rule
bites on PBS (`token ∩ user`).
## TLS
**SHA-256 leaf-cert pinning** against the self-signed host cert. **No insecure default**, ever. The
pin is the raw leaf-DER sha — the SAN is never checked, so a cert rotation changes the pin and the
agent must be re-pinned.
## The destructive path — never the direct call
| Do not | Because | Use |
|---|---|---|
| `Client.DestroyLXC` / `Vzdump` / `SetConfig` ad-hoc | skips classification, signature, per-guest serialization, crash recovery | `reconcile.Engine` paths / `RunSignedJob`; queue via `Queue.Submit` |
| add a method to `proxmox.Privileged` | breaks the 3-exception root-CLI fence (`routing_test.go`) | `proxmox.Runner` + a new sudoers `Cmnd_Alias` + `validate.go`-style checks |
| treat `ListLXC` output as "guests we own" | audit A1 — pre-v0.62.0 the stale-lock reaper did exactly this, contained only by the pool-scoped token | intersect with `Client.Pool` membership (`staleLockController.Guests()`); **fail safe on read failure** |
Full trap table: `REUSE.md` §3. Every guest joins the `felhom` pool — `VM.Audit` comes from the
`/pool` grant, not from a per-guest ACL.
<!--
The fence is not stylistic. It is what makes this component auditable: two types, one of which can
only speak HTTP and one of which can only shell out, with a test asserting neither crosses. A single
convenience method on Privileged that also makes an HTTP call would end that property silently.
-->
+49
View File
@@ -0,0 +1,49 @@
---
paths: ["internal/storage/**", "internal/escrow/**"]
---
# Storage and escrow — format safety and zero-knowledge recovery
`internal/storage/` is the storage observer, durable IDs, role/claim classifiers, `SudoHostOps` and
the watchdog. `internal/escrow/` is the PBS-key escrow with its zero-knowledge recovery code.
> **Overlap note:** `health-checks.md` also matches `internal/storage/**`. Deliberate — both rules
> apply there and both load.
## Never format the device you inspected
**AGENT-001 is a TOCTOU:** acting on the caller's `req.Device` (or any remembered `/dev` path) after
inspection lets `/dev` re-enumeration retarget the node to a **different physical disk**. Format the
**re-resolved** device — `Server.reresolveWipe` / `reresolveBlank`.
**Never exec raw `mkfs.*`** (including `Binaries.MkfsExt4`/`MkfsXfs`): sudoers no longer allowlists
raw mkfs, and going direct bypasses the claim filter and the wrapper's re-checks. Use
`SudoHostOps.Format`, which routes through `felhom-mkfs-guarded`.
## The two durable-ID schemes refuse each other
They are not interchangeable, and each returns a `binding_mismatch` for the other's scheme:
| Purpose | Scheme | Resolver |
|---|---|---|
| wipe confirmation | `byid:` / `byuuid:` | `ResolveDurableDevice`, `DiskInfo.WipeDurableID` |
| enrolled-storage remount | `uuid:` | `ResolveStorageDevice` |
Using `DiskInfo.DurableID` (a `uuid:`) as a wipe-confirmation id is F20-BUG2.
## Drive data is never taken by force
Plain `umount` only — **never `-l`, never `-f`**, and never any format operation under
`/mnt/felhom-drives`.
## Escrow is zero-knowledge, and a fetch failure is not a wrong code
The server holds no client key; a no-key restore fails with `missing key`. **A fetch failure must
never be reported as a wrong recovery code** — that told a customer their correct code was bad, in
hundredths of a second, when checking a code actually takes about one. Distinguish "we could not
reach the store" from "the code did not match", always.
<!--
The escrow recovery-code "flake" was a REAL defect, not a flake. "Known flake, re-run" needs evidence
before it is said out loud — that phrase cost this project a real finding once.
-->
+35
View File
@@ -29,6 +29,41 @@ root=$(git rev-parse --show-toplevel 2>/dev/null) || {
}
cd "$root" || exit 1
# ── WORKSPACE-ROOT ASSERTION (2026-08-05, R-204 rider) ───────────────────────────────────────────
# Refuse a push from a clone outside the felhom workspace.
#
# WHY THIS IS A HOOK AND NOT A LINE IN A DOCUMENT: the workspace root is ALREADY written down, in
# documentation/runbooks/workspace-CLAUDE.md and in the workspace-root CLAUDE.md ("stay inside it"),
# and work drifted into a home directory anyway. A rule that has failed once as a reminder is not
# fixed by writing it down again — it has to be asserted where it can bite.
#
# A PUSH IS THE RIGHT TRIGGER, deliberately: throwaway clones under /tmp for probes and red-proofs
# never push, so nothing legitimate breaks. Reads and builds elsewhere stay unaffected.
#
# Symlinks are resolved on BOTH sides before comparison, so a symlinked path neither falsely passes
# nor falsely fails. If the workspace root does not exist on this machine the check is SKIPPED, not
# failed — this hook must not brick a legitimate clone on a different host.
#
# The only bypass is the documented `git push --no-verify`, whose use is already reportable.
FELHOM_WORKSPACE_ROOT=/mnt/5_hdd/felhom.eu
if [ -d "$FELHOM_WORKSPACE_ROOT" ]; then
ws_real=$(cd "$FELHOM_WORKSPACE_ROOT" 2>/dev/null && pwd -P) || ws_real=""
root_real=$(pwd -P) || root_real=""
if [ -n "$ws_real" ] && [ -n "$root_real" ]; then
case "$root_real/" in
"$ws_real"/*) : ;; # inside the workspace — proceed
*)
echo "pre-push: PUSH REFUSED - this clone is OUTSIDE the felhom workspace." >&2
echo " clone: $root_real" >&2
echo " expected: under $ws_real (repos live in $ws_real/git/<repo>)" >&2
echo " Work in the workspace clone, or bypass with 'git push --no-verify'" >&2
echo " and state that you did in the session report." >&2
exit 1
;;
esac
fi
fi
if ! command -v python3 >/dev/null 2>&1; then
echo "pre-push: FAIL - python3 not found, so the gates CANNOT run. This is a failure, never a" >&2
echo " pass by default. Install python3, or push with --no-verify and say so." >&2
+559
View File
@@ -1,3 +1,560 @@
## v0.133.0 — a restore-test can never fill a box's disk; leftovers retried on a timer (2026-09-24, R-672, R-673)
> **RELEASED 2026-09-24** by `scripts/release-agent.sh` — tag `v0.133.0` (`9bdb4da`), sha256 `3aa303452b8c6be58573d00af01a0ab4a0d97e4f885ffecd0144a18ac24e69b6`, verified by an independent anonymous download. **NOT delivered** (needs an operator-signed `agent_update` job per box, R-530) and **NOT vouched**.
**MinAgent impact:** none required by any controller. Hub **v0.124.0** makes a thin pool CRITICAL at 90 %
(data or metadata) and keys the storage-fill alarm per pool per 6 h; an older hub still raises its generic
90/95 % storage-fill events from the same report.
- **R-672 — the space preflight** (`reconcile/restoretest_space.go`, provider `internal/restorespace`). Before
anything is journaled or created, a restore-test needs free data ≥ restored × 1.2 + 5 GiB on its target
(`backup.restore_test_space_factor` / `backup.restore_test_space_reserve_gib`) and room in a thin pool's
metadata. `restored` is the UNCOMPRESSED size — the vzdump log's "Total bytes written" (a file-backed
archive) or the PBS snapshot size; the archive FILE is never used (9201: 6.9 GB file, 22.6 GB restore). The
target moves OFF the tested guest's own pool when another storage is eligible (active, `rootdir`, and the
agent holds Datastore.AllocateSpace there) and fits. Anything unknown refuses. A refusal is reported to the
hub as the test's result (`pass=false`, `skipped=true`, "skipped: not enough space on …") — never a pass,
never dropped. The archive config is read once (it was read twice). Measured case: demo-hp 2026-09-24 —
the restore-test that filled `local-lvm` would now be refused (needs 32.1 GB, 23.2 GB free).
- **R-672 — a failed scratch teardown is retried every 10 minutes** (`Engine.RetryScratchTeardown`, the
daemon's janitor), not only by Recover at agent start; never a vmid a running test owns; after 3 failed
tries the operator is told through a failed restore-test record naming the scratch guest.
- **R-672 — a thin pool crossing 90 % requests an immediate host report** (`Observer.SetThinHighTrigger`, the
storage watchdog's read path, every few seconds; re-armed below 85 %), so the hub's alarm sees it in
seconds, not at the next 15-minute report.
- **R-673 — the stale-lock sweep runs on the same 10-minute timer** (it ran only at start: a stale
`snapshot-delete` lock blocked 9201's whole-box backups for five hours), holding the one-heavy-operation
gate so no agent backup can start between its "no vzdump running" check and its unlock.
- Red-proofs: six, each seen failing (REPORT).
## v0.132.0 — a controller that dies slowly is reported, not just restarted (2026-09-17, R-539)
> **RELEASED 2026-09-17** by `scripts/release-agent.sh` — tag `v0.132.0`, sha256 `4afe815749a41b327ebe4a98a4557ad2acffb1fa71740a3a8b8003473b835321`. **Vouched 2026-09-17** for Day-0 installs on the operator's word, together with golden **0.246.0** (the hub's R-120 gate required the newer golden first). Delivered to demo-hp and the N100 by an operator-signed `agent_update` job each (operator ruling 3 of 2026-09-16), not by a floor.
**MinAgent impact:** none required by any controller. Hub **v0.117.0** turns the new fields into
`controller_slow_crashloop`; an older hub ignores them.
- **R-539 (operator ruling 3 of 2026-09-16) — the slow crash-loop counter.** The 3-restarts-in-15-minutes
brake cannot see a controller that dies every 20 minutes (measured 2026-09-16, R-531: four restarts,
none accumulating, only an `info` event that mails nobody). Beside it, unchanged, a second counter:
restarts the supervisor performed in the last **24 hours**; at the **fifth**, the heartbeat's
`controller_supervisor` stanza sets `slow_crashloop_since` (and `slow_crashloop: true`,
`restarts_24h`). The hub mails on that timestamp MOVING; it moves **at most once per 24 hours**. It
does **not** stop restarting — the fast brake remains the only brake.
- **Persisted per guest** at `/var/lib/felhom-agent/guests/<vmid>/controller-restarts-24h.json`
(tmp + rename, 0600), so an agent restart or a host reboot does not reset it. The fast record stays
in memory; the reason it does (a persisted "give up" could outlive the fix) does not apply to a
counter that only warns. Unreadable or corrupt → WARN and a clean start, never a blocked supervisor.
- **Deliberate kills count.** The supervisor cannot tell an operator's `docker kill` from a crash
(measured 2026-09-15); a controller killed five times a day is worth a line either way.
- The supervisor's startup line now prints `slow_crashloop_max=5 slow_crashloop_window=24h0m0s`.
**Red-proofs, each seen failing:** five restarts 20 minutes apart raise it
(`TestControllerSupervisor_SlowCrashloop` — fails with no counter; fails again with the once-per-24-hours
guard removed, "the operator would be mailed per restart"); the counter survives an agent restart
(`…SlowCounterSurvivesAgentRestart` — fails with the save removed, `Restarts24h:1`); restarts seven hours
apart never raise it (the negative control). Wire shape extended with `restarts_24h` and `slow_crashloop`.
## v0.131.0 — a dead controller comes back by itself; the backup status speaks per tier (2026-09-15, R-523 / R-517 / R-518)
> **RELEASED 2026-09-15** by `scripts/release-agent.sh` — tag `v0.131.0`, sha256 `1118b552f7e775fbde9544c7764ede7e6046e0a7db16ae8494d07a18e3c2ac9c`. Delivered to boxes by the controller v0.243.0 floor (declared MinAgent), not by hand.
- **R-523 (P1) — the in-guest controller supervisor.** BIGNIGHT F9: `docker kill felhom-controller`
left the household's dashboard on 502 for 33 minutes, because nothing watched the container.
Measured first (2026-09-15, Docker 29.8.0, `evidence-p1fixes-2026-09-15/A1`): after `docker kill`,
BOTH `--restart unless-stopped` and `--restart always` leave the container `exited (137)` after
60 s — a policy change alone is not a fix. New `internal/localapi/controllersupervisor.go`: every
30 s, for each felhom-pool guest the agent provisioned (`<guests>/<vmid>/bootstrap` exists) that is
running, it reads `docker inspect -f {{.State.Status}} felhom-controller`; on the SECOND consecutive
not-running (or absent) observation it runs `systemctl restart felhom-controller-bootstrap.service`
inside the guest — the swap's own restart, over the same GuestExecutor and the same two sudoers
grants. No new privilege. Guards, each pinned by a test: not during a controller swap (the swap's
in-flight flag); not when parked (`touch /var/lib/felhom-agent/guests/<vmid>/controller-parked` on
the HOST); not on a stopped, locked or vzdump-busy guest; not on an unknown docker answer; not on a
guest the agent did not provision; and **no thrash** — 3 restarts in 15 minutes stop the restarts
for 30 minutes. The record rides the host report as `controller_supervisor` (additive,
`omitempty`); hub v0.114.0 mints `controller_restarted_by_agent` (info) and `controller_crashloop`
(error), both operator-only. Red-proofs: without the restart call, the kill test fails at
"restarts=0"; without the backoff block, the crash-loop test fails at "restarted 10 times".
- **Golden script** (`configs/build-golden.sh`): the controller runs `--restart always` (covers a
Docker daemon restart after a manual stop — nothing more). **No golden baked here** (R-468); existing
boxes keep `unless-stopped` until their next golden and are covered by the supervisor.
- **R-517 (P1) — `GET /backup/status` speaks per tier.** The untargeted response gains `tiers[]`:
per tier the newest SUCCESSFUL backup (`last_success`, from the record, or from the tier's storage
after an agent restart — `last_success_source: storage`), the last attempt kept apart
(`last_attempt {started_at, success, error}`), and whether the tier's storage exists (`storage:
present|absent|unknown`). `GET /backup/tiers` gains the same `storage` field (R-518's cheap half: the
controller skips an absent tier). `unknown` is never `absent` — a storage view that cannot be read
must not skip a backup. Additive; the untargeted `.backup` keeps its meaning (pinned). Red-proof:
filling `last_success` from the newest ATTEMPT fails at "pbs tier reports a failed attempt as its
last success".
## the decoy sweep — can this gate be fooled by a label? (2026-09-01, R-421) — NOT A RELEASE
**No product code, no version bump, no image, no golden.** A scripts change is not a release.
Four times in one week a gate turned out to match a NAME instead of the thing it named — R-410 (a
`mkdir` turned the release gate green), R-400 (seven debug controls answering nothing), R-378 (a
status word inside a sentence), R-419 (a phrase inside prose, including prose saying the marker was
absent). **All four found by accident.** The gates enforce everything else here and were the one part
nothing had checked.
**All 29 gate scripts read and decoyed. 16 were fooled.** 10 fixed here, 4 left with rows
(R-422..R-425), 6 could not be given a plausible decoy and are named (R-426 group d).
**The largest single cause was mundane:** eight gates set their SCOPE with `os.listdir` (one level).
Green and correct today; blind the moment anyone adds `templates/partials/`. `mojibake` and
`docker-v` already used `os.walk`, caught the identical planted file, and are the control that
proves the cause was the listing rather than the decoy.
Full survey table, and the five decoys withdrawn as illegitimate (mine, named):
`documentation/audits/AUDIT-gate-decoys-2026-09-01.md`.
**In this repo:** no gate changed, and that is the result. `release-complete` was decoyed and is
SOUND — the sweep's attempt (a non-version heading on top of `CHANGELOG.md`) was WITHDRAWN as
illegitimate, because `HEAD_RE.search` scans the whole file and still names `v0.130.0`. The three
shared gates are covered from `felhom.eu`; the remaining two are named in the decoy-coverage
exemption list (R-426) as UNTESTED, not as sound.
## v0.130.0 — the agent was the one leaking connections onto the off-site box (2026-08-20, R-344)
> **RELEASED 2026-08-20**, on the operator's word, after the fix was proved on both boxes.
> `sha256 a56a92a7bd68f5b46736eaec4806c3d26c16ccb35118c4ac0e3d8094eaefabc3`, 14,141,158 bytes,
> tag `v0.130.0` at `7569f34`. Reproducible: a rebuild with `-trimpath -buildvcs=false` matches the
> published artifact byte for byte (R-186's property, checked rather than assumed).
>
> The heading read `## UNRELEASED — v0.130.0 candidate` until this point, deliberately: while the fix
> was hand-installed on `demo-hp` only, publishing would have pushed it onto `demo-felhom` through
> self-update and destroyed the control the proof rested on. **`release-complete` convicted on the
> release heading and was right to** — the answer was to stop claiming a release, not to bypass the
> gate. See R-347.
>
> **The fleet runs these exact bytes.** Both demo boxes were first given a hand build made during the
> proof — same source, same version string, **different bytes** (`256e0829…`), because
> `release-agent.sh` builds with `-trimpath -buildvcs=false` and a hand build does not. Nothing would
> have corrected that: the boxes already reported `0.130.0`, so self-update saw the vouched version as
> installed and would have done nothing, forever. Both were reinstalled from the **downloaded package**
> and now report `a56a92a7…`. Filed as **R-349**, because every prove-then-publish train hits it.
**What was measured, before anything was changed.** Between 2026-08-18 09:51:22Z and 2026-08-20
08:02:13Z, ep0's PBS proxy accumulated **388 established connections** — 194 from each demo box, on a
proxy whose descriptor ceiling is 65536 and whose runway at that rate was ~323 days. The connections
were held open on **both** sides: ep0 showed 388 while the two boxes showed 194 + 194, at two separate
instants, with the same source ports on each side, and **not one closed in a 31-minute window**.
`ss -tnp` on the boxes named the holder: **`felhom-agent`**, 194 of 194 on each, one PID.
**`pvestatd` and `proxmox-backup-client` made 162,404 requests in that window and leaked zero.** They
are 99.5% of the traffic to that endpoint and 0% of the leak. The agent made 811 requests — of which
387 were `GET .../snapshots` — and leaked 388 sockets. One per call, within one.
**The defect, and it is two things compounding.**
- `internal/pbs/client.go` built its transport as a composite literal:
`&http.Transport{TLSClientConfig: tlsCfg}`. That takes **`IdleConnTimeout` zero, which does not mean
"use a sane default" — it means retain idle keep-alive connections FOREVER.**
`http.DefaultTransport` sets 90s; hand-rolling the transport (which every client here must do,
because they all pin TLS) silently discards it.
- `pbsTargetsFromPVE` (`cmd/felhom-agent/main.go`) builds **a fresh `pbs.Client` every cycle**, as its
own doc comment says, and drops the previous one. An abandoned `http.Transport` does **not** close
its connections — it becomes unreachable while its `persistConn` read-loop goroutine keeps the
socket alive. So each cycle stranded exactly one connection that nothing could ever close.
The cadences reconcile without fitting: a 900s live-snapshot collect (184.7 cycles in the window) plus
a 6-hour verify loop (7.7) predicts 192.4 per box against **194 observed**.
**The fix is one field, restored to the standard library's own value.** New leaf package
`internal/httpx` owns `DefaultIdleConnTimeout = 90 * time.Second` — 90s because that is what
`http.DefaultTransport` uses, so there is nothing invented here to justify or tune — and
`NewTransport(tlsCfg, idleConnTimeout)`, which returns a **fresh** transport (never shared: each caller
pins a different endpoint) and treats a zero or negative timeout as **use the default, never "no
timeout"**. `pbs.Config` gains an `IdleConnTimeout` field that production leaves unset; only tests set
it, to avoid a 90-second wait.
**`internal/hub/client.go` and `internal/proxmox/client.go` carried the identical missing default and
were corrected in the same pass — but neither contributed to the ep0 leak, and this entry must not be
read as three leaks having been found.** Both are built **once per process**, so they held one idle
connection for the life of the daemon rather than accumulating, and neither talks to ep0:8007.
**Tests, and what they deliberately do not assert.** `internal/pbs/client_leak_test.go` counts
connections **server-side** and models what `pbsTargetsFromPVE` actually does — build a client, use it
once, drop it on the floor — then asserts the connections go away. It does not assert `err == nil` and
it does not assert that some field holds some value; both were true of the leaking code.
- **Red-proof 1 (the fix):** removing `IdleConnTimeout` from `NewTransport` fails the test with
*"after 5s the server still holds 5 open connection(s), want 0 (5 dialled in total)"* — the count is
in the message, so the failure cannot be mistaken for a timeout with another cause. Reverted.
- **Red-proof 2 (the fix that would be worse than the bug):** setting `DisableKeepAlives: true` also
makes the leak vanish — by dialling fresh for every request, which on a box polling ~40,000 times a
day is strictly worse than what we started with. **The leak test PASSES under that mutation**;
`TestPBSClient_KeepAliveStillReuses` is what catches it, failing with *"3 sequential requests over 3
connection(s), want 1"*. Reverted.
**What this release does NOT do.** It does not reduce the poll rate (**R-336 stays open, but re-scoped
— it was never the cause of this leak**), it does not refactor `pbsTargetsFromPVE` to cache or reuse
clients (a one-line default restores the standard behaviour; a lifecycle refactor adds
cache-invalidation questions for no measurable gain), and it adds no `CloseIdleConnections` call.
**One sentence in this entry was written before the deploy and was WRONG, and it is corrected here
rather than quietly edited.** It read: *"does not clear the 388 descriptors already stuck on ep0 —
those persist until that proxy restarts."* **Measured: they clear the moment the AGENT restarts.**
Replacing the binary on `demo-hp` released exactly its 199 descriptors within one second
(415 → 216 fd), and replacing it on `demo-felhom` released the remaining 203 (**220 → 17 fd in under
two seconds**). **17 is precisely ep0's `t0` baseline** of 2026-08-18 09:51:22Z. ep0 was read-only
throughout and its proxy PID never changed. The accumulated leak was never ep0's to hold on to — it
was held on both sides, and closing either side ends it.
## v0.129.0 — a correct code for an earlier package stops being called wrong (2026-08-12, R-311)
**The measurement this fixes.** On 2026-08-12 a recovery code that provably opens a RETAINED package
— unsealed by hand, and it restored planted files byte-identical from a store the box itself could no
longer open — was answered by this agent with *"the recovery code did not open the sealed bundle"*.
The code was correct. Nothing had ever tried the retained packages, so the engine could not tell a
correct-but-earlier code from a mistype, and the screen said so out loud: a true sentence about our
own incuriosity, read by the customer as a statement about their code.
**`OffsiteKeyRecoverer` gains an optional `FetchRetained`.** It is consulted ONLY after the current
package has refused, so the ordinary recovery pays nothing for it and cannot fail because of it. When
one of the retained packages opens, the recoverer returns `ErrCodeOpensRetained` wrapped in a
`RetainedOpenedError` carrying the supersession date — no material, no code, no password.
**The local API answers 422** ("your code is correct, it belongs to an EARLIER sealed package") — a
FIFTH status added to the R-224 switch, not a restructuring of it. 422 rather than 400 because the
request was well-formed AND the credential valid; a 400 would put it in the same bucket as a mistype,
which is the defect.
**Fail-safe in every direction.** A nil fetcher, a hub too old to have the route (404 is a clean
"none"), a transport failure, a malformed package: each leaves the original refusal standing,
unchanged. The worst outcome of this feature breaking is the behaviour we had before it existed.
Attempts are bounded (`MaxRetainedTried`, default 6) because each unwrap is ~1 s of scrypt by design
and an unbounded loop would turn one wrong code into a minutes-long hang.
**New hub client call:** `FetchRetainedIdentityEscrow` → `GET /api/v1/hosts/<id>/escrow/retained`
(hub >= v0.103.0), self-scoped by the same per-host key.
Seven tests with REAL age crypto, because the two situations are indistinguishable AT THE UNWRAP and a
faked unwrap would prove nothing about what was broken. Red-proof, asserted applied: removing the
retained lookup returns the fail-closed wrong-code error — **the lie comes back, in those words.**
---
### Gates only — 2026-08-09 (no release, no version bump, no binary published)
**Two guards, both owed since the 2026-08-09 install outage (R-273/R-287). Nothing that runs on a
customer's box changed; `scripts/` only, and the agent stays v0.128.0.**
- **`scripts/retention-policy.json` — THE retention number, in one file.** The registry stopped
serving `felhom-agent` 0.120.0 and older while `check-published-versions.py` demanded that every
tag still be downloadable. Both rules are sensible; together they are impossible, and CI went red
at a commit whose own run had been green the day before. The check now **reads the number from
this file** and bounds its assertion to the newest N generic versions.
**What CI no longer covers, said plainly rather than left to be discovered:** a released version
older than the retention window is **no longer asserted downloadable**. Its git tag and its configs
are still asserted — only the binary's presence is dropped. The check **prints exactly which
versions it stopped covering** on every run, so the narrowing cannot become permanent by accident.
**The number is an OBSERVED state, not a located ruling** — see the file's own header and R-287.
A missing or unreadable policy file is **INCONCLUSIVE (exit 2), never silently unbounded.**
- **`scripts/check-release-complete.py` — the tag half, as a machine.** Asserts that the version at
the head of `CHANGELOG.md` is tagged, that the tag points into this history, and that its package
is published. `release-agent.sh` already warned about this in as many words and the step was still
missed on 2026-08-08, which is why this is a gate and not a reminder. Legs 1–2 need no network and
therefore run in `--fast`, so the pre-push hook catches a missing tag at the earliest moment.
Registered in `agent_gates.py`; red-proved by pointing the CHANGELOG head at an unreleased
v0.129.0 — both legs convicted and each named its fix command.
## v0.128.0 — the escrow seed is asserted every tick, not remembered once (2026-08-08, R-221)
**A rebuilt box could not run the escrow ceremony at all, and there was no way forward from inside
the product.** The preflight refuses on `escrow.pbs_storage_id`; the pbsdr bridge writes that key;
and it wrote it from exactly one place — `finishConverged`, reached only on the paths that actually
converge.
**The two things live in different places and die at different times.** The convergence marker is
host-side (`<agent-state>/pbsdr/marker.json`). The key it seeds is in `agent.json`, which
`step_agent_config` renders from `base = {}` unless an explicit `--preserve-from` is passed
(`felhom.eu/scripts/felhom-host-install.sh:2396`, the render at `:2449`, the `O_TRUNC` write at
`:2579`; the flag at `:1246`, defaulting empty at `:256`) — **and the render never writes an
`escrow` section at all.** So a rebuild keeps the marker and takes the key: same descriptor, same
hash, early return, and the seed never runs again into a config that no longer has it.
**Fixed by asserting rather than remembering.** `Apply` now re-asserts the seed *before* the
idempotent early return. `seedEscrowStorageID` is unchanged and still never clobbers a different
existing value — an operator's own choice outranks the descriptor's, with a warning naming both.
**The early return is KEPT.** It exists so a converged box does not re-run Proxmox operations every
60 s, and `TestSeedReasserted_OnConvergedTick_WithZeroProxmoxCalls` asserts **zero** recorded runner
calls on that tick — so a "fix" that simply deleted the return would fail. Cost of the re-assert: one
small file read plus a JSON parse per tick, no exec, no network, and an early return once the value
matches.
**A seed failure can never un-converge the box:** Warn plus a message on the published status,
exactly as `finishConverged` does it — no marker write, no state change. Pinned by
`TestSeedReassertFailure_DoesNotUnconverge`.
Tests drive the real `Apply` with a real temp-dir `agent.json` and a call-recording runner; calling
`seedEscrowStorageID` directly cannot see the early return, which IS the defect. The production
wiring (`pbsdr.NewManager(..., cfg.SourcePath, ...)`) is asserted by walking `main.go`'s **AST**, not
by `strings.Contains`, which a commented-out call also satisfies.
Red-proofs: removing the new call makes Scenario A fail on today's tree; removing the early return
makes the zero-Proxmox-calls assertion fail.
## (no version bump) — a comment that claimed the hub reads a field it has no field for (2026-08-08, R-260)
Comment-only; no behaviour, no wire change, nothing to rebuild.
`HostReport.SelfUpdatePending` / `SelfUpdatePendingVersion` carried the sentence *"The hub reads an
absent field as pending=false, the correct default."* **The hub has no field for either**, so it reads
nothing — present or absent — and `encoding/json` discards them on arrival. The sentence described an
intent rather than the code and read as settled for long enough that a class sweep had to find it.
The emission is correct and stays: the agent reports the truth, and the fault is entirely in the
receiving. The missing consumer is tracked as **R-264** (OPEN), and
`felhom.eu/scripts/wire_contract_gate.py` now refuses any NEW field of this shape while recording the
existing ones as explicit, reasoned allowlist entries rather than silence.
## v0.127.0 — a mount Felhom itself made is not "something else" (2026-08-06, R-220)
**After a rebuild the customer's own drives could not be re-attached, and the refusal named an action
they could not perform.** `GET /api/disks/candidates` returned `initialize: [], attach: []` while both
drives sat there, and the deploy refused with *"choose an attached drive from the list"* — a list that
was empty. Measured live **three times**: CAMPAIGN-11 Phase 1, and twice on the R-201 re-walk.
**The mechanism.** Enrolment mounts a drive **twice** — at the managed `/mnt/felhom-drives/<name>` and
at the raw `/mnt/<name>` it creates on the host. **The host survives a guest rebuild; the controller's
registry does not.** So `classifyClaim` saw a mount outside the managed prefix and correctly concluded
"claimed by something else" — about Felhom's own mount.
**The fix is CORROBORATED, not a widened prefix.** A mountpoint outside `/mnt/felhom-drives` is
forgiven **only when the same device is ALSO mounted under the managed path** — a pairing that only
Felhom's own enrolment produces. A disk another system is using, at `/srv/data` or `/media/x` or even
`/mnt/someone-elses-disk`, has no such counterpart and **is still refused**. That fence has its own
test, and its red-proof shows an over-wide fix offering `/mnt/someone-elses-disk` for formatting.
**Read from `/proc/mounts`, deliberately.** The lsblk invocation is pinned **verbatim** in
`configs/felhom-agent.sudoers` (`lsblk -J -o NAME,FSTYPE,PTTYPE,MOUNTPOINT /dev/*`), so switching it to
the plural `MOUNTPOINTS` would have meant shipping a sudoers change with the binary — a far larger
blast radius than this finding warrants. `/proc/mounts` is world-readable: **no sudo, no new allowlisted
command, no config change.**
**Fail-safe:** an unreadable mount table corroborates **nothing**, so the device classifies exactly as
it did before this change — refused. "We could not corroborate" must never read as "it is ours".
Tests: `claim_r220_test.go` — the own-drive case, the foreign-mount fence over four paths, and the
corroboration itself (both mounts required; a lone raw mount vouches for nothing; another device's
managed mount does not vouch for this one; an unreadable table corroborates nothing).
**Red-proofs:** removing the exemption refuses the customer's own drive again
(*"device is mounted at /mnt/adatok (sdb)"*); over-widening it to any `/mnt/*` path breaks the fence.
## docs — CLAUDE.md becomes a core plus path-scoped rules (2026-08-06, R-229 leg (b)) — no version bump
**Documentation only. No Go changed, nothing built, nothing deployed.** `go build`/`vet`/`test` green
and unchanged.
**175 -> 99 effective lines** (207 -> 103 raw, 14,093 -> 6,267 bytes). The release/publish-train
section was the largest block and the `felhom-build-deploy` skill already carries the procedure, so
the core points at it instead of restating a table that drifts from the script. The package layout
went the same way as the controller's: `REUSE.md` and the tree are its home, and the per-package
annotations that were doing real work moved into the rule file for the area they describe rather than
being deleted.
**New:** `.claude/rules/{proxmox,localapi,backup,storage}.md`, all `paths:`-scoped, all <=46 effective
lines, joining the existing `health-checks.md`.
**Kept in the core deliberately** — it is the only part re-injected after `/compact`: the root-CLI
fence and its three named exceptions (breaching it is how this component stops being auditable), the
destructive-op gate, the prove-ownership rule from audit A1, the gate entry point, the F9
live-validation fence, trunk-based with its revert-and-report escape hatch, and the end-of-session
checklist.
**Glob overlap, stated rather than silently resolved:** `health-checks.md` matches
`internal/{localapi,guesthook}/**` and `internal/storage/**`, which `localapi.md` and `storage.md`
also match. Both rules load in those directories and neither supersedes the other; each new file says
so in its own text so a reader who sees two rules fire is not left guessing which wins.
## docs — the "CI is still owed" claim was stale; corrected (2026-08-06, R-229 part 2) — no version bump
**One sentence, no code.** This file asserted that continuous integration was still owed
(`felhom.eu` `OPEN-ITEMS.md` R-168). **R-168 was CLOSED on 2026-08-02** — a Gitea Actions runner
re-runs each repo's gate entry point on every push and emails the operator on failure. Found while
confirming this session's own push by run ID, which is the check that caught it.
The same stale sentence was in four instruction files across all four repos and is corrected in all
four. In `felhom-agent/CLAUDE.md` it **contradicted the same file's release section**, which already
said R-168 mails the failure — a contradiction inside one instruction file, which is the exact class
the R-229 work exists to find.
## docs — expired and contradictory blocks removed from CLAUDE.md (2026-08-06, R-229) — no version bump
**Documentation and gate registration only. No Go changed, nothing built, nothing deployed.**
Surgical corrections; the file was deliberately **not** restructured (that is deferred, R-229).
- **Deleted the expired TEMPORARY block.** It read *"felhom-pve is at a remote site (until
~2026-08-02) … Delete this block on return"* and was still being read as current fact on
**2026-08-06**, four days past its own deadline — while `felhom-controller/CLAUDE.md` asserted the
opposite. The location-independence fact worth keeping (`localapi` binds `169.254.253.1:8443` on
`vmbr9` since the R-50 island migration) moved to an HTML comment.
- **Every component version literal is gone** from effective text, including
`felhom-agent --version → 0.115.0` and the `go.mod` Go directive. Versions change several times a
day; ask the hub's `/hosts` + `/configs` or the box.
- The drill-VM claim and the host addresses now point at `documentation/operations/nodes.md`, which
already stated both correctly. **This file's drill-VM claim was the correct one** — confirmed by
`qm list` on demo-hp.
- The R-115/R-188/R-186 release **narratives** moved to an HTML comment and to the
`felhom-build-deploy` skill; the **directives** stayed (never hand-roll the build; the
build → tag → publish → push order; reproducible `-trimpath -buildvcs=false`).
- The health-check block-I/O rule became `.claude/rules/health-checks.md`, scoped to the five
packages where health checks are written. It had been duplicated from `felhom.eu/CLAUDE.md` *with a
note explaining that that file does not load in an agent-only session* — correct reasoning, made
obsolete by path-scoped rules.
`agent_gates.py` now registers **`instructions`** (shared, `felhom.eu/scripts/`, never copied).
Full accounting: `felhom.eu/documentation/audits/LEDGER-instruction-trim-2026-08-06.md`.
## v0.126.0 — a fetch failure is not a wrong recovery code (2026-08-06, R-224)
**A hub the agent could not reach was reported to the customer as a bad recovery code.** Measured live
on 2026-08-05 (CAMPAIGN-11 F3): with the hub REJECTed at the appliance's firewall and a **correct,
current** recovery code, the customer was told their code did not open their package — **in 0.0556 s**,
against ~1.0 s for a genuine unseal. No unseal was attempted. F4 produced the same message in 0.0299 s
with this agent stopped.
**The discriminator existed here the whole time and this boundary threw it away.** `recover.go` fails
at four distinguishable points; the local-api handler had cases for two of them and a `default` that
answered *"the recovery code did not open the sealed bundle, or the bundle could not be fetched"* —
one sentence for two situations, only one of which is the customer's doing.
**The fix is a value, not a log line.** `escrow.ErrBundleFetch` joins the fetch leg's error, and the
handler routes it to **502** with its own words: *"the sealed recovery bundle could not be fetched from
the hub — the recovery code was NOT used and nothing was written."* 502 rather than 4xx because the
request was not bad; an upstream dependency failed. The `default` now carries **only** the fail-closed
wrong-code case and says so without the "or".
Four situations, four statuses — **502** fetch failed · **400** the bundle was fetched and refused the
code · **404** the hub holds no bundle · **409** the bundle predates the repository-password field.
The controller classifies on the STATUS and must never parse these sentences.
⚠ **A GREEN TEST NAMED THIS DEFECT AND DID NOT PREVENT IT.**
`TestRecoverOffsiteRepoPassword_FetchErrorIsDistinct` has said since v0.125.0 that *"the operator must
not be sent to re-read their recovery code because the hub was unreachable"* — and it passed
throughout, because it asserted this package's error **string** one layer below where the merge
happened, and a string is something no caller can branch on. It now asserts the sentinel, and its
consequence-level twin asserts the STATUS at the boundary the customer's message is derived from.
**Prefer the test that asserts the consequence over the one that asserts the mechanism.**
Tests: `recover_test.go` (fetch classifies as `ErrBundleFetch`; a wrong code does **not**; an absent
blob keeps its own identity) and `localapi/escrow_recover_class_test.go` (each situation's status, and
a standalone assertion that fetch-failure and wrong-code never share one). **Red-proofs:** removing the
`%w` join fails the sentinel test; deleting the handler case makes both answer `400` with the
wrong-code sentence — the exact pre-fix code, and the exact defect CAMPAIGN-11 measured.
## v0.125.0 — the agent opens the sealed bundle and returns one field (2026-08-04, R-199 links 7–8)
**Link 7 had one production caller and it was a `--selftest`.** `UnwrapIdentityBundle` has existed
since slice 10D.1 and the only thing that ever called it was `runSelftestIdentityConsume`, reading the
recovery code from an environment variable by hand. **Link 8 did not exist at all:** that selftest
writes the whole bundle JSON to a file, and its success message named `tunnel_token + pbs_token` —
an enumeration that was accurate when written and became a MISSTATEMENT the moment v0.77.0 sealed the
offsite repository password into the same bundle. Anyone reading that output would conclude the
repository password was not there. It now names what THIS bundle actually carried and what it did not.
**`POST /escrow/recover-offsite-password`** on the pinned local API: the controller supplies the
customer's recovery code, the agent fetches this host's own sealed blob from the hub
(`hub.Client.FetchIdentityEscrow`, hub ≥ v0.94.0, self-scoped by the per-host key), unseals it, and
returns **only the offsite restic repository password** plus its sha256.
**Only that field, on purpose.** The bundle also carries the tunnel token, the PBS token and the WG
private key. The controller is a trust tier down and needs none of them; returning them would widen
the blast radius of a controller compromise for nothing. Narrowing costs nothing now and is not
recoverable later.
**Why the agent and not the controller:** `age` is an agent runtime dependency and is deliberately
absent from the controller image; the blob is a host-scoped object whose only writer is this agent
under the per-host key, so the read is that write's mirror.
**R's handling is the tightest rule in this release.** It arrives in the request body over the pinned
channel, is held in memory for one call, is cleared on the success path AND every failure path, is
never written to disk, never an argument in a process list, never logged at any level including
inside an error, and is never echoed. `UnwrapIdentity` already stages only the blob and the recovered
plaintext in a temp dir it removes; a test redirects TMPDIR and asserts **the tree is empty
afterwards** — emptiness rather than a content scan, because a content scan is defeated by a later
call overwriting the leaked file, which is exactly how the first version of that test passed its own
red-proof while R sat on disk.
Three outcomes are distinct rather than one generic failure: no blob (404 — no ceremony has run), a
bundle that opens but predates the field (409 — a pre-fork-4 blob, which cannot be retro-fitted), and
a code that does not open it (400 — fail-closed at the KDF, nothing written). Sending an operator to
re-check a correctly typed recovery code because the hub was unreachable is the mistake this avoids.
**The wiring is asserted by an AST walk**, not a `strings.Contains`: `main` → `runDaemon` →
`buildLocalAPIServer`, where an `escrow.OffsiteKeyRecoverer` is constructed and passed as
`localapi.Options.EscrowRecovery`, and its fetcher calls the DAEMON's own hub client (the self-scoping
that makes cross-host retrieval impossible is a property of which key is used). This project's
built-but-never-wired count is six and links 6–7 were two of them; the fix must not become the seventh.
## v0.124.1 — the repair record must survive the probe that did NOT feed the hub (2026-08-04, R-190)
**v0.124.0's transition record did not reach the hub, and the live run is what showed it.** The
capability reported degraded for "one cycle" — meaning the probe call that performed the repair. But
`probeAll` is invoked **independently** by the periodic self-check log and by the collector building a
host-report. On the demo box the repairing call was the log's (`09:39:34`, journal shows
`GRANT WAS MISSING AND HAS BEEN SELF-REPAIRED` and `degraded=1`), and the host-report built three
seconds later found the grant present and sent **`ok`**. The agent's journal had the record; the hub
had nothing; the operator would have learned nothing.
That is the exact silence R-190 is about, re-created inside its own mitigation — and every unit test
passed while it was true.
**The fix is a latch on TIME rather than on call count.** A confirmed repair is reported for
`storeGrantRepairReportWindow` (20 minutes), which comfortably exceeds the 900 s host-report interval,
so at least one report must carry the transition. It clears on its own — a permanently degraded
capability would be its own false alarm — and it is per tier.
**Two hollow tests were caught and fixed on the way**, both the same shape this repo keeps finding: a
test asserting a value it constructed itself, and a test asserting the latch HELPER rather than the
path that consumes it — whose red-proof duly passed. The decisions now live in
`storeGrantHealthyVerdict` and `storeGrantRepairedVerdict`, and the tests call those.
## v0.124.0 — a lost storage grant repairs itself, and says that it was lost (2026-08-04, R-190)
**R-190 is a grant that demonstrably worked at 04:44 on 2026-08-03 and was gone by 09:24** — with a
host reinstall, logged `pveum` activity and cluster-log entries all ruled out by measurement. The
cause is still open. The resilience does not have to wait for it.
**Everything needed already existed and had only ever been called once.** The root wrapper
(`felhom-backup-target-apply grant <id>`), its sudoers vector (`grant *`, any storage id, confirmed
not assumed), and the exact command were all in place — and the `grant` verb had only ever run at
storage CREATION. That is the *built but never wired* shape, in a verb rather than a seam, and it is
this project's seventh instance.
**What v0.124.0 does:** when the store-grant probe finds the grant absent on a tier the box depends
on, it runs that wrapper and **re-reads once** to confirm — the pbsdr R-22 self-grant shape, including
its restraint: one attempt, one confirmation, and anything still wrong stays loudly wrong.
**THE RECORD IS THE POINT, AND IT IS THE HALF R-190 IS ACTUALLY ABOUT.** A repair that leaves only
`ok` behind destroys the only evidence a permission vanished, so a recurring loss becomes undetectable
forever — strictly worse than the fault it fixes. So a confirmed repair reports **DEGRADED for exactly
one cycle**, with the explanation in `Feature`:
```
backup tier felhom-backup: the agent's storage grant was MISSING and has been AUTOMATICALLY
RESTORED — the tier works now, but a permission that vanished on its own needs investigating (R-190)
```
**Nothing new was built to carry it.** The hub's existing ok→degraded→ok edge is the channel — it
alerts and e-mails on the first edge and logs the recovery on the next cycle, so one loss produces
exactly one alert pair. No wire change, no hub change, no new event type. `Feature` carries the text
because that is the field the hub interpolates into the operator's e-mail; `Reason` does not travel.
**Bounded (Scenario F):** one attempt per tier per hour, in memory. A storage can be unreadable for
reasons an ACL cannot fix, and a re-grant on every report cycle is a repair loop wearing a fix's
clothes. An agent restart re-arms it, which is correct — a restart is exactly when a box should
re-check what it depends on.
**A failed repair never masks the fault:** the capability stays degraded with the failure in its
reason, and a repair that "succeeded" but did not survive the re-read is reported as needing a human.
## v0.123.0 — a tier the box cannot READ now says so (2026-08-03, R-185)
**The missing permission is one command. The silence was the defect.** On demo-felhom the agent's PVE
@@ -5037,3 +5594,5 @@ client, signing, or storage/backup orchestration yet (later slices).
read-only `--selftest` against the demo host with TLS fingerprint pinning.
- The 16-privilege `FelhomAgent` role + privsep token (role on **both** user and
token) is provisioned out-of-band; the agent only consumes the token.
<!-- R-421 sweep: this repo cites R-421; the row landed in felhom.eu 2d88776. -->
+88 -193
View File
@@ -1,216 +1,111 @@
# CLAUDE.md — `felhom-agent`
> Loads when Claude Code touches this repo. Stable orientation only — **current state lives in
> `CONTEXT.md` and the top of `CHANGELOG.md`**, never here. Cross-repo orientation: workspace-root
> `/mnt/5_hdd/felhom.eu/git/CLAUDE.md`.
> Stable orientation only — **current state lives in `CONTEXT.md` and the top of `CHANGELOG.md`**,
> never here. Cross-repo conventions (artifact taxonomy, access, clean-tree gate, secrets,
> CHANGELOG/REPORT): workspace-root `/mnt/5_hdd/felhom.eu/git/CLAUDE.md`. Path-scoped detail:
> `.claude/rules/`.
## What this repo is
`felhom-agent` is the operator-tier **host agent** that runs on each Proxmox host and owns **all**
Proxmox interaction: provision/restore guests, host storage, backup/restore orchestration, the hub
control loop, and a narrow per-guest local API. It is the **most privilege-sensitive** component.
The operator-tier **host agent**, one per Proxmox host, owning **all** Proxmox interaction:
provision/restore guests, host storage, backup/restore orchestration, the hub control loop, and a
narrow per-guest local API. It is the **most privilege-sensitive component in the system**.
- Renamed former `proxmox-controller` repo.
- **Distinct from `felhom-controller`** — that is the *in-guest* controller (Docker-only, no Proxmox
creds). Do not confuse them.
- Control plane, not data plane: if the agent dies, apps keep serving; only management degrades.
- Renamed from `proxmox-controller`.
- **Distinct from `felhom-controller`** — that is the *in-guest* controller, Docker-only, holding no
Proxmox credentials. Do not confuse them.
- **Control plane, not data plane:** if the agent dies, apps keep serving; only management degrades.
- Pure Go stdlib + `golang.org/x/crypto`. No web frameworks.
## Read before writing code
## Doing X → read Y
- **`REUSE.md`** — canonical helpers, format-safety guards, traps, seams. Check it first; update it
in the same commit that changes a shared helper or pattern.
- `CONTEXT.md` (current state + open threads) and the top `CHANGELOG.md` entry (authoritative history).
- Design doc: `felhom.eu/documentation/architecture/03-host-agent.md` (locked). Platform facts:
`felhom.eu/documentation/proxmox-platform.md` + `tests/phase{0,1-2,3,4}-findings.md`.
| Doing | Read |
|---|---|
| writing any new code | `REUSE.md` — helpers, format-safety guards, traps, seams |
| needing current state / open threads | `CONTEXT.md` + the top `CHANGELOG.md` entry |
| Proxmox, reconcile or signed jobs | loads itself: `.claude/rules/proxmox.md` |
| local API, authz or guest hooks | loads itself: `.claude/rules/localapi.md` |
| backup, PBS or DR | loads itself: `.claude/rules/backup.md` |
| storage or escrow | loads itself: `.claude/rules/storage.md` |
| writing a health check | loads itself: `.claude/rules/health-checks.md` |
| **release, build, publish, deploy, verify a version** | the **`felhom-build-deploy`** skill — **never hand-roll it** |
| writing or reviewing a test, fixing a bug | the **`felhom-testing`** skill |
| host addresses, break-glass, node facts | `felhom.eu/documentation/operations/nodes.md` — never restate them |
| which box may I break | `felhom.eu/documentation/runbooks/target-selection.md` |
| what version is live anywhere | ask the hub (`/hosts`, `/configs`) or the box — **never a doc** |
| the authoritative design | `felhom.eu/documentation/architecture/03-host-agent.md` (locked) |
## Layout (verified against the tree)
## The root-CLI fence — API-first, exactly three exceptions
```
cmd/felhom-agent/ main + flags + --selftest modes + the daemon entry
cmd/felhom-opsign/ offline operator signing CLI (SSHSIG)
internal/authz/ operator signed-op verifier (SSHSIG) + durable FileNonceStore
internal/backup/ vzdump backup runner + restore-test scheduler + report store
internal/capability/ live sudo-policy capability probe (degradation visibility)
internal/config/ JSON config + FELHOM_AGENT_* env overlay; secrets redacted (Redacted())
internal/desired/ hub desired-state syncer (envelope observer)
internal/escrow/ PBS-key escrow (zero-knowledge recovery code)
internal/guesthook/ pre-start self-heal hookscript install
internal/hub/ daemon: HostReport collector + Bearer client + resilient Loop
internal/lanresolver/ split-horizon DNS on guest IP change (dnsmasq RESTART, not reload)
internal/localapi/ per-guest local API: token store, disks/format, guest binds, controller swap,
stale-lock recovery, pinned self-signed leaf
internal/log/ slog setup
internal/pbs/ PBS-API client (fingerprint-pinned) + verify maintenance loop
internal/provision/ guest bootstrap back-half (token mint → bootstrap.json → pct bind)
internal/proxmox/ API-first Client + fenced root-CLI Privileged + UPID WaitTask
internal/reconcile/ reconcile engine + reversibility gate + op journal + crash recovery
internal/signedjobs/ operator-signed destructive executors (wipe, decommission)
internal/storage/ storage observer + durable ids + role/claim classifiers + SudoHostOps + watchdog
```
## Build / run
- Module `gitea.dooplex.hu/admin/felhom-agent`; binary `felhom-agent` (`cmd/felhom-agent/`).
- **Pure Go stdlib + `golang.org/x/crypto` only** — no web frameworks. `go.mod` directive go 1.25.0;
DooPlex (192.168.0.180, where CC runs) has the Go toolchain and is on the same LAN as the demo
host — build and run live tests locally.
- Version via `-ldflags "-X main.version=<v>"`; `--version` flag. Bump on meaningful changes + CHANGELOG entry.
- **Full build/deploy/publish runbook: use the `felhom-build-deploy` skill.** Summary:
> **Clean-tree gate before any build:** `git status --porcelain` must be empty and
> `git rev-parse HEAD` must equal `git rev-parse origin/main` in the repo being built. An unpushed
> change does not exist — never build a dirty or unpushed tree. The `git pull` in the build step
> stays (it is a no-op when you work in this tree, and load-bearing if anything was pushed from
> elsewhere).
> **RELEASING IS ONE COMMAND, AND IT PUBLISHES (R-115).** There used to be a raw `go build` line
> here and a *separate* "Publish" row, so publishing was a step someone had to remember — and it was
> **forgotten three times in five days**, the last leaving agent v0.120.0 deployed on both demo hosts
> and undownloadable, where a documented-path reinstall would have silently downgraded them while
> reporting success. Do not hand-roll the build: the script also creates the `v<version>` git TAG
> that `felhom-host-install.sh` fetches this version's sixteen config files from (R-183), and it
> verifies by an **independent download** rather than trusting the publish step's own output.
> `scripts/publish-agent.sh` still exists and is still correct — the release script CALLS it rather
> than reimplementing it.
>
> **THE ORDER IS build → tag LOCALLY → publish → push tag, and each step protects something (R-188,
> R-186).** The tag is created before the publish so the build and the tag describe the same commit;
> it is *pushed* after, because the push is what wakes CI (`on: [push]`) and a tag visible before its
> package makes the published-versions gate correctly fail a correct release — it did, on roughly
> every second release, and R-168 sends that failure to you by mail. The invariant the old order
> protected is asserted directly instead: the gate now also refuses a **published version with no
> tag**. If the push fails after a successful publish the script says so loudly and prints the
> one-line recovery; if the *publish* fails it removes the local-only tag so a retry is clean.
>
> **A RELEASED BINARY IS INDEPENDENTLY VERIFIABLE (R-186).** The build uses `-trimpath
> -buildvcs=false` so the same source produces the same bytes whether or not the tag exists yet —
> before this, a rebuild could not reproduce the sha you were vouching. To check any published
> version yourself:
>
> ```bash
> V=0.122.0
> git checkout "v$V" && go build -trimpath -buildvcs=false -ldflags "-X main.version=$V" \
> -o /tmp/felhom-agent-check ./cmd/felhom-agent
> sha256sum /tmp/felhom-agent-check
> curl -fsSL "https://gitea.dooplex.hu/api/packages/admin/generic/felhom-agent/$V/felhom-agent" | sha256sum
> ```
>
> The two hashes must match. `publish-agent.sh`'s fallback build uses the **same** flags — it used to
> force `CGO_ENABLED=0` and produce a 74 KB-smaller binary for the same version; if either build line
> ever changes, change both or one version name means two binaries again.
| Step | Where | One-liner |
|---|---|---|
| **Release** (build + tag + publish + verify) | DooPlex (local) | `GITEA_USER=admin GITEA_TOKEN=<tok> scripts/release-agent.sh <ver>` — refuses a dirty/unpushed tree and refuses to re-release an existing version |
| Copy | local → felhom-pve | `scp /tmp/felhom-agent-<v> felhom-pve:/tmp/` (one hop) |
| Deploy | felhom-pve | backup `.bak-<old>` → `install -m0755` → `systemctl restart felhom-agent` (non-root `felhom-agent` user, config `/etc/felhom-agent/agent.json`) |
| Ship configs | felhom-pve | sudoers (`/etc/sudoers.d/felhom-agent`) + guarded-mkfs wrapper WITH the binary when `configs/` changed |
| **Verify** (anyone, any time) | anywhere with the repo + Go | `git checkout v<ver> && go build -trimpath -buildvcs=false -ldflags "-X main.version=<ver>" -o /tmp/a ./cmd/felhom-agent && sha256sum /tmp/a` — must equal `curl -fsSL <pkg-url> \| sha256sum` |
| **Vouch** | hub operator UI | Configs → Day-0 artifacts. **Deliberately NOT automated** — vouching is what points machines at a version, and it stays your act (prove-then-vouch) |
| Verify | felhom-pve | `felhom-agent --version` + journal (clean ReassertGuestBinds, no capability degradation) |
## Proxmox model (the load-bearing rules)
This is in the core because breaching it is how this component stops being auditable.
- **API-first** via a scoped `FelhomAgent` token. Raw root-CLI is **fenced to exactly 3 exceptions**:
keyctl `pct create` (golden image), USB mount/fstab, SMART/sensors. `Client` never shells out;
`Privileged` never makes HTTP calls (asserted by `routing_test.go`). Keep that fence.
- **Every mutating op is async** → returns a UPID → `WaitTask` asserts `exitstatus == "OK"`. A 200 on
the POST is **not** success; authorization can fail at task execution.
- **TLS:** SHA-256 leaf-cert pinning (self-signed host cert). No insecure default.
- **Privsep token gotcha:** a `--privsep 1` token's rights = intersection of the backing user's perms
AND the token's ACLs — the role must be granted on **both**, or every call 403s.
- Destructive ops go through the reconcile gate / signed-jobs path — never call `Client.DestroyLXC`/
`Vzdump`/`SetConfig` ad-hoc (REUSE.md §3).
keyctl `pct create` (golden image), USB mount/fstab, SMART/sensors.
- **`Client` never shells out; `Privileged` never makes HTTP calls** — asserted by `routing_test.go`.
Adding a method to `proxmox.Privileged` breaks the fence; use `proxmox.Runner` plus a new sudoers
`Cmnd_Alias` and `validate.go`-style checks (`REUSE.md` §3).
- **Destructive ops go through the reconcile gate / signed-jobs path.** Never call
`Client.DestroyLXC` / `Vzdump` / `SetConfig` ad-hoc — that skips classification, signature,
per-guest serialization and crash recovery.
- **Ownership must be PROVEN, never assumed.** A raw `ListLXC` list is not "guests the agent owns";
intersect with `Client.Pool` membership and fail safe on a read failure (audit A1).
## Demo host (for live tests)
## Gates — ONE entry point
Node **`demo-felhom`**, API `https://192.168.0.162:8006`. SSH alias `felhom-pve` (root@pam) —
available to CC as plain `ssh felhom-pve`. A **second demo node `demo-hp`** (HP t740, node name
`felhom-host`, `ssh demo-hp` — no baked key; break-glass root via hub `host_recovery/demo-hp-bb76ea` +
`sshpass`) is the **designated drill+build VM host** per the 2026-07-25 operator ruling, and that ruling
is **realized** — it hosts drill VM `300` (`drill-r50`), so **start there**, not on DooPlex. (The
historical golden-bake `drill.qcow2` still lives on DooPlex and is a bake fixture, not a drill target.)
**Which box is safe to break, and what may be done to each:
`felhom.eu/documentation/runbooks/target-selection.md`** — read it before any destructive test. Both
nodes + the break-glass recipe: `felhom.eu/documentation/operations/nodes.md`. The agent pins the served leaf cert — verify the
fingerprint still matches before a live run. Selftest modes (run locally on DooPlex, pointed at the
demo API): `--selftest[=read|task|hub|storage|backup|restore-test|pbs-verify]`; no flag = the daemon.
**Run `python3 scripts/agent_gates.py` from the repo root after ANY change here.** It runs this
repo's gates — `reuse_refs_check` and `instructions_gate`, both the **shared** copies in
`felhom.eu/scripts/`, never copied into this repo (a copy recreates the drift they detect; an absent
sibling clone FAILS). `--fast` selects the gates touching no network and no container runtime; today
that is all of them. **A missing gate is a FAILURE, never a skip.**
> **TEMPORARY — felhom-pve is at a remote site (until ~2026-08-02).** The home-LAN literal
> `192.168.0.162` is NOT reachable from DooPlex for the duration. Access via Tailscale:
> felhom-pve = 100.70.170.35; the `Host felhom-pve` entry in `~/.ssh/config` on DooPlex already
> points there (the direct-LAN path stays available as `Host felhom-pve-lan`). Delete this block on
> return. All documented `ssh felhom-pve` / `pct exec` workflows are unchanged. Path is **direct**
> (not DERP), ~37 ms rtt per hop. At the remote site the host is on **DHCP**; re-check its address
> rather than trusting one written here (`ip -br addr show vmbr0` — it read `192.168.0.162/24` on
> 2026-07-30, and `felhom-pve-lan` from DooPlex is still `No route to host`). Details + findings:
> `felhom.eu/documentation/audits/AUDIT-vacation-remote-ops-2026-07-20.md`
>
> **The "agent does not run at the remote site" warning this block used to carry is RETRACTED
> (2026-07-30) — it was true before R-50 and is false now.** `localapi` no longer binds a LAN literal:
> since the R-50 island migration (2026-07-25) it binds `169.254.253.1:8443` on `vmbr9`, which is
> location-independent by design, and `proxmox.endpoint` is `https://127.0.0.1:8006`. Verified live:
> `systemctl is-active felhom-agent` → `active`, `felhom-agent --version` → 0.115.0, and the per-guest
> local API answered `GET /disks` over the island. No config edit and no Viktor GO are outstanding.
**The pre-push hook** (`.githooks/pre-push`) runs it with `--fast` and refuses a failing push. It is
**per-clone** — switch it on once with `git config core.hooksPath .githooks`, and a manual run WARNS
when this clone is unarmed. `git push --no-verify` bypasses it deliberately; **say so in the session
report when you use it** — CI re-runs the same entry point on every push and **emails the operator on
failure**, so a bypass is noticed even though it is not blocked (R-168, CLOSED 2026-08-02).
> **Legacy: Windows workstation.** Until 2026-07-19 CC ran on Windows 11; `pct` commands over SSH
> needed `export MSYS_NO_PATHCONV=1`, and every remote command used
> `SSH=/c/Windows/System32/OpenSSH/ssh.exe`. Agent deploy was a two-hop copy via the Windows box
> (`cygpath -w` for the local scp path; CRLF hazard on config files).
<!--
WHY ONE ENTRY POINT (2026-08-02, R-29): a census of all gates across the four repos found every check
a CLAUDE.md names was passing, and two of the four nobody is told to run were failing. This repo was
the extreme case — nothing ran against it at all, and 90 cited paths were checked by no one.
-->
## Live validation — the fence
Exercise the **SERVER-SIDE PIPELINE** a real user triggers, end-to-end. **The forbidden shortcut is
BYPASSING it** — the F9 episode was a raw guest-attach with hand-set state, and it proved nothing.
`claude-in-chrome` is NOT available on DooPlex. Invoking the exact endpoint the UI invokes is an
acceptable proxy — **say which method was used**. Low-level mechanism tests where the direct call IS
the mechanism are exempt.
## Conventions
### Trunk-based — no branches
All shippable work commits **directly to `main`**; `main` equals what is deployed.
- Report-only artifacts (audits, findings, fixspecs) → `felhom.eu/documentation/` (`audits/`, `backlog/`).
- Risky/supervised fixes are spec'd, then implemented **during the supervised session, on `main`**.
- Unattended escape hatch: if a fix can't be cleanly verified/shipped, revert + report — never park on a branch.
> **In every repository where you make a change, update both files in that repo:**
> - **`CHANGELOG.md`** — cumulative log, newest on top.
> - **`REPORT.md`** — **overwrite** with the most recent implementation/validation summary only.
>
> **Never write secrets** into any committed file — reference them as "stored out-of-band".
- Code quality: verify generated code for bugs/edge cases; add debug logging; **ask rather than
guess** when you'd otherwise invent input/output.
- **A health check issues no block I/O** — no `statfs`, no `getdents`, no read, write or `fsync`, **not
even behind a timeout**. Liveness is decided from `/proc` and kernel state. The full rule + the
measurement lives in `felhom.eu/CLAUDE.md` "Code quality rules"; it is repeated here because health
checks are written in THIS repo and that file does not load in an agent-only session. R-117 spike §6.3.
- Update `REUSE.md` if you added/changed/deprecated a shared helper or pattern (same commit).
- **Run `python3 scripts/agent_gates.py` from the repo root after ANY change in this repo.** It is
the ONE entry point for this repo's gates. Today it runs one — `reuse_refs_check` over this
repo's `REUSE.md` — and it exists at one gate on purpose: a census on 2026-08-02 found that every
check a `CLAUDE.md` names was passing and two of the four nobody is told to run were failing, and
this repo was the extreme case, with nothing running against it at all and 90 cited paths checked
by no one. It grows when the agent grows a second check. `--fast` selects the gates that touch no
network and no container runtime; today that is all of them. A missing gate is a FAILURE, never a
skip. **The shared `reuse_refs_check.py` lives in `felhom.eu/scripts/` and is never copied here**
— a copy would recreate the drift it detects; an absent sibling clone FAILS the gate.
**The pre-push hook** (`.githooks/pre-push`) runs it with `--fast` and refuses a failing push. It
is per-clone — switch it on once with `git config core.hooksPath .githooks`, and a manual run
WARNS when this clone is unarmed. `git push --no-verify` bypasses it deliberately; **say so in the
session report when you use it.** Both facts are why CI is still owed (`OPEN-ITEMS.md` R-168).
- Testing doctrine (non-hollow tests, red-proofs, seams): use the `felhom-testing` skill.
- **Logging**: the slog logger fans out to journald (configured level) + the always-DEBUG `applog.Ring`
(remote pulls) — English, keys-never-values, durations on outcomes; full rules in
- **Trunk-based — no branches.** All shippable work commits directly to `main`; `main` equals what is
deployed. Report-only artifacts (audits, findings, fixspecs) go to `felhom.eu/documentation/`.
- **Unattended escape hatch:** if a fix cannot be cleanly verified and shipped, **revert and report**
— never park it on a branch.
- **Logging**: the slog logger fans out to journald (configured level) plus the always-DEBUG
`applog.Ring` (remote pulls). English, keys-never-values, durations on outcomes. Full rules:
`felhom.eu/documentation/runbooks/logging-conventions.md`.
- Update `REUSE.md` in the same commit that adds, changes or deprecates a shared helper or pattern.
### Live validation
## End-of-session checklist
Exercise the SERVER-SIDE PIPELINE a real user triggers, end-to-end. The forbidden shortcut is
BYPASSING it (the F9 episode: raw guest-attach + hand-set state). Invoking the exact endpoint the UI
invokes is an acceptable proxy when a browser isn't available — say which method was used. Low-level
mechanism tests where the direct call IS the mechanism are exempt.
- **`CHANGELOG.md`** (cumulative, newest on top) and **`REPORT.md`** (overwritten with this run only)
— in every repo touched.
- **`CONTEXT.md`** — decisions, state, what is next.
- **`REUSE.md`** — if a shared helper or pattern moved.
- **A finding goes in `felhom.eu/documentation/backlog/OPEN-ITEMS.md` first**, never only in a report
or an audit.
- **Confirm your own last push's CI run went green, by run ID** — CI mails on failure, which is a PUSH
signal; this is the PULL check that catches a lost or unread mail. An unchecked green is an
assumption, not an observation.
## Workflow & artifacts
- Implement **`TASK.md` / `TASK-*.md`** specs (when placed as `TASK.md` or told to), then push +
CHANGELOG + REPORT.md.
- **`RUNBOOK-*.md`** — an operational procedure. CC executes the steps it has access and capability
for, including live validation on the demo Proxmox host (CC has root@felhom-pve SSH + the
felhom-agent token). Mark a step HUMAN only when it genuinely needs physical presence, a real-world
decision, or credentials CC truly lacks. Judgment still applies: confirm before irreversible ops on
real customer data — demo scratch guests are fair game.
**A gate ships with a decoy test that has been seen to fail (R-421).** A decoy is the LABEL without
the FACT — a directory with the right name and no bake log, a note whose prose mentions the marker it
lacks. `scripts/decoy_coverage_gate.py` refuses a new gate that has neither a decoy nor a named
exemption carrying its row. The four shapes, the 2026-09-01 sweep that fooled 16 of 29 gates, and the
decoys withdrawn as illegitimate: `documentation/audits/AUDIT-gate-decoys-2026-09-01.md` and
`felhom-controller/.claude/rules/gates.md`. **Scope is a fact too** — prefer `os.walk` over
`os.listdir`, and a glob over a hand-maintained list.
+71
View File
@@ -1,10 +1,81 @@
# CONTEXT — felhom-agent working state
> **2026-09-24 — v0.133.0 RELEASED, NOT DELIVERED (R-672, R-673).** Restore-test space preflight
> (`reconcile/restoretest_space.go`, `internal/restorespace`): uncompressed size from the vzdump log / PBS size,
> × 1.2 + 5 GiB, thin metadata, off the tested guest's pool, unknown refuses, reported as `skipped` non-pass.
> Janitor (`cmd/felhom-agent/janitor.go`) every 10 min: `Engine.RetryScratchTeardown` + stale-lock sweep under
> the heavy-op gate. Thin pool ≥ 90 % → immediate report (hub v0.124.0 alarms). **Operator rulings 2026-09-24
> (evening):** the scheduled restore-test is OFF on both demo hosts (`backup.restore_test_eval_interval_seconds:
> -1` — note: 0 means the 6 h DEFAULT, only a negative disables) until v0.133.0 is delivered there; saved configs
> `/etc/felhom-agent/agent.json.pre-r672`. demo-hp 9201 was repaired (stop, fsck, start; two Redis AOF tails cut).
> Snapshot of the current state + open threads. Authoritative history lives in `CHANGELOG.md` (top
> entry = current); the end-of-task detail lives in `REPORT.md`.
## R-199 (v0.125.0) — links 6–8 of the recovery chain, assembled and walked
`POST /escrow/recover-offsite-password` (pinned local API, `withGuest`): the controller supplies the
customer's recovery code, the agent fetches THIS host's own sealed blob from the hub
(`hub.Client.FetchIdentityEscrow` → `GET /hosts/{id}/escrow`, hub >= v0.94.0, self-scoped by the
per-host key), unseals it via `escrow.OffsiteKeyRecoverer`, and returns **only** the offsite restic
repository password plus its sha256.
**Rules that must not erode:**
- **Only that field.** Not the tunnel token, not the PBS token, not the WG key — the controller is a
trust tier down and needs none of them. Narrowing cost nothing and is not recoverable later.
- **The unseal stays in the agent.** `age` is an agent runtime dependency (`/usr/bin/age` — hardcoded,
no config override; 1.2.1 on demo-felhom) and is deliberately absent from the controller image.
- **R:** in memory for one call, cleared on the success path AND every failure path, never on disk,
never in argv, never logged at any level including inside an error, never echoed. Verified live: 0
log lines, 0 files, 0 leftover `felhom-idesc-*` dirs, with a positive control proving the search worked.
- **Three distinct outcomes**, not one generic failure: no blob (404), a bundle that opens but predates
the field (409 — pre-fork-4, cannot be retro-fitted), a code that does not open it (400 — fail-closed
at age's KDF, nothing written).
- **The wiring is pinned by an AST walk** (`cmd/felhom-agent/escrow_recover_wiring_test.go`):
`main` → `runDaemon` → `buildLocalAPIServer`, an `escrow.OffsiteKeyRecoverer` constructed there, the
`Options.EscrowRecovery` field present, and the fetcher calling the DAEMON's own `hubClient` (the
self-scoping that makes cross-host retrieval impossible is a property of WHICH key is used).
Links 6 and 7 were two of this project's six built-but-never-wired instances.
**Proven live on demo-felhom 2026-08-04:** recovered sha256 == on-disk sha256 == the hub's stored hash.
A wrong code five minutes earlier failed closed. **The chain stops at link 8** — nothing installs a
recovered password, reopens a repository, or restores a file.
**§8.6, fixed while here:** `runSelftestIdentityConsume`'s success line used to recite
"tunnel_token + pbs_token", which became a misstatement when v0.77.0 sealed the repository password
into the same bundle — anyone reading it would conclude the password was not there. It now names what
THIS bundle carried and what it did not.
## Current
- **2026-08-03 — v0.123.0 (R-185): a tier the box cannot READ now says so.** The agent's token had
`FelhomAgentStore` on `local`, `local-lvm`, `felhom-pbs` and **not** on `felhom-backup` — the
storage both demo boxes configure as `local_backup_target`. That storage answered `{"data":[]}`
through the token while root listed three archives, and `pickForThisRun` skipped it as *"no settled
archive yet"* — **which is what a brand-new tier reports**, so the host tier was never
restore-testable and nothing said so.
- **The permission question is asked directly**, because unlike the listing it has a definite
answer: `Client.Permissions` reads `/access/permissions?path=/storage/<target>` **as the agent's
own token**, and `storeGrantStatuses` emits one `capability.Status` per configured tier. It
composes AROUND the sudo prober, the way `poolReadStatus` already does — an API read does not
belong inside a sudo-policy probe. `Status`'s wire shape is untouched, so the hub's critical
degraded alert applies with **no hub change**.
- **MEASURED FIRST, and the obvious reading is wrong:** an ungranted path answers neither empty nor
403 — it carries the privileges INHERITED from the box-wide `/` grant
(`Sys.Audit, SDN.Use, Datastore.Audit`). Checking path-presence, or `Datastore.Audit`, reports a
blinded storage HEALTHY. The probe tests **`Datastore.AllocateSpace`**; re-measure before ever
changing that constant (`storeGrantRequiredPriv`, red-proved).
- **The probed set comes from `BackupTiers()`, never a fixed list** — a hardcoded probe list is the
defect reproduced inside the fix. Critical, EXCEPT the `local` fallback target (reported, but it
does not page). It never consults content, so it cannot alarm on a newborn tier; it never reports
ok when it could not ask.
- **LIVE:** degraded observed on the still-blind box (hub emailed `agent_capability_degraded`) →
grant applied on **both** demo boxes → token lists 3 and 4 archives → `ok=70 total=70 degraded=0`
and `degraded → ok` at the hub → **the host tier became a due-check candidate for the first time**,
correctly picking the 08-02 archive (08-03 had not settled 24 h).
- **The installer's real defect was NOT `PVE_STORAGES`** — see `felhom.eu` CONTEXT S-22: Case A
grants, the Scenario-F reuse arm did not. Fixed in installer **1.24.0** with a gate.
- **2026-08-03 — v0.122.0 (R-189 · R-188 · R-186): three signals that lied about their own work.**
None touches data; all three cost attention, which every other signal depends on.
- **R-189 — a passing restore-test no longer vanishes on a restart.** `restore_tests[]` came only
+23 -275
View File
@@ -1,281 +1,29 @@
# REPORT — R-189 · R-188 · R-186: three ways the signals lied about themselves
# REPORT — agent v0.133.0: a restore-test can never fill a box's disk (2026-09-24, R-672, R-673)
**Date:** 2026-08-03 · **Repo:** `felhom-agent` **v0.121.1 → v0.122.0** (`7581f81`) · released,
published, verified by an independent download **and by rebuilding it**, deployed to demo-felhom.
`felhom.eu`: register + docs only, **no hub change and no hub bump** — the hub already reads
`restore_tests[]`; the defect was that the agent stopped sending them.
Full record: `felhom.eu/documentation/audits/r672-2026-09-24/README.md`.
---
## Shipped (released, not delivered)
Tag `v0.133.0` (`9bdb4da`), package sha256 `3aa30345…e69b6`, verified by anonymous download. Delivery needs the
operator's signed `agent_update` job per box.
## 1. Baselines, re-read on arrival
- **Space preflight** before anything is created: free ≥ restored × 1.2 + 5 GiB, `restored` UNCOMPRESSED (vzdump
log "Total bytes written" / PBS snapshot size), thin metadata with room, off the tested guest's pool when
another eligible storage fits, unknown refuses, reported as a non-pass (`skipped`).
- **Scratch teardown retried every 10 min**; operator told once after 3 failed tries.
- **Thin pool ≥ 90 %** → immediate host report (hub v0.124.0 alarms `storage_fill_critical`, per pool per 6 h).
- **R-673:** the stale-lock sweep on the same timer, under the one-heavy-op gate.
| Repo | `main` @ commit | Version | Matched §1? |
|---|---|---|---|
| `felhom-agent` | `3d0a1d615d11` | `v0.121.1` | **yes** |
| `felhom.eu` | `c9a3e48b2106` | hub `v0.91.1` | **yes** |
## Red-proofs (each seen failing; files in `felhom.eu/documentation/audits/r672-2026-09-24/redproofs/`)
1. preflight removed → the 2026-09-24 restore issued again; 2. archive FILE size used → "the compressed file size
was used"; 3. timer pass a no-op → "the leaked scratch was not destroyed by the timer"; 4. sweep without the gate →
"the sweep ran while a backup held the gate"; 5. no 90 % edge → "0 report requests, want 1"; 6. every skip dropped →
"a space refusal never reached the host report". Full suite `go test ./...` green; `agent_gates.py` OK.
Highest register ID in use was **R-189**; no new IDs were needed — all three rows already existed.
(Grep confirmed R-190+ free, in case one had been.)
## Live (demo-hp, the on-demand `--selftest=restore-test`, cadence OFF)
(a) factor 10 → refused: needs 215.5 GiB, has 22.1 GiB. (b) normal margin → refused: restoring 21.1 GiB needs 30.3 GiB,
has 22.1 GiB — correct: no full restore-test fits demo-hp under 80 % pool use. (c) a forced teardown failure could not
run live (it needs a scratch guest, which (b) shows cannot be created within the 80 % rule) — unit tests only. Pool
58.99 % before and after every run; nothing created.
## 2. Scenario H — the reproducibility measurement (Part 3, done first on purpose)
**Before**, one commit, same source, same toolchain, same ldflags — the only difference is whether the
tag existed when the build ran:
| build | sha256 | size | embedded module version |
|---|---|---|---|
| default flags, **no tag yet** | `18f4a495…` | 14 085 464 B | `v0.121.2-0.20260803133646-3d0a1d61` |
| default flags, **tagged** | `4a38f394…` | 14 085 440 B | `v0.121.99` |
| `-trimpath -buildvcs=false`, either way | `7ffcdf1d…` | 14 064 574 B | *(none)* |
**After, on the real release (v0.122.0) — the three values §15.2 asks for:**
| artifact | sha256 | size |
|---|---|---|
| published, downloaded from Gitea | `d5f294e56c1ef59055e8e87fb9135aa477632dbbc56a5d4bff46bbd0466c1edf` | 14 076 649 B |
| rebuild at the tag, #1 | `d5f294e56c1ef59055e8e87fb9135aa477632dbbc56a5d4bff46bbd0466c1edf` | 14 076 649 B |
| rebuild at the tag, #2 | `d5f294e56c1ef59055e8e87fb9135aa477632dbbc56a5d4bff46bbd0466c1edf` | 14 076 649 B |
**All three identical.** The property is removed-cause, not sequenced-around: `-buildvcs=false` drops
a stamp nothing reads (no `ReadBuildInfo` caller, verified by grep), and the version still comes from
the explicit `-X main.version` ldflag. `-trimpath` additionally makes a rebuild from a different
checkout directory match.
**A second discrepancy fell out of the measurement and is fixed with it:** `publish-agent.sh`'s
fallback build forced `CGO_ENABLED=0` and produced **13 990 236 B** against the release path's
**14 064 574 B** — a 74 KB difference, i.e. one version name meaning two binaries depending on which
entry point ran. Both paths now use identical flags, each commented with a pointer to the other.
## 3. Scenario E — a correct release no longer emails a failure (Part 2)
**Only the tag PUSH moved.** The order is now build → tag **locally** → publish → push tag. The tag is
still created before anything is published, so the build and the tag describe the same commit; it
becomes *visible* — to CI (`on: [push]`) and to any `raw/tag/…` fetch — only once the package is
downloadable.
**The old ordering's invariant is asserted directly rather than arranged for.**
`check-published-versions.py` now carries two invariants: every tag has an installable package (as
before) **and no published version is missing its tag** (new). The second is a **bounded probe** —
the frontier past the newest tag, where a failed tag push leaves an orphan, plus patch gaps — and it
**prints its probe set on every run**, because a check whose coverage is invisible reads as a
guarantee it is not making. The package listing api was **re-measured**, not assumed: `401` without a
token, so absence still cannot be enumerated, and the script says so in its own output.
**Live result — this release is the test:**
| release | CI runs on the release commit | outcome |
|---|---|---|
| v0.121.0 (yesterday) | #12 / #13 | one **green**, one **red** |
| v0.121.1 (yesterday) | #17 / #18 | one **red**, one **green** |
| **v0.122.0 (this one)** | **#21 (task id 96) / #22 (task id 97)** | **both green** |
**Failure modes made loud rather than tidy:** a publish that succeeds followed by a tag push that
fails now dies printing `git push origin v<ver>` — the local tag is already there, so recovery is one
line — and a publish that *fails* deletes the local-only tag so a retry is clean instead of colliding
with step 2's re-release guard.
## 4. Scenarios F and G — both directions demonstrated, then cleaned up
**G — a published version with no tag must FAIL.** A real fixture: version **0.121.2** (the frontier,
exactly where a failed tag push lands) published to the live registry with no tag.
```
0.121.2 next patch after the newest tag PUBLISHED — NO TAG
check-published-versions: 1 PUBLISHED VERSION(S) WITH NO TAG
v0.121.2 is downloadable at … but has no git tag.
git push origin v0.121.2
EXIT=1
```
**Red-proof (observed):** with the fixture still live, replacing the probe set with `[]` produced
`ALL RELEASED VERSIONS INSTALLABLE, AND NONE UNTAGGED`, **exit 0** — a green run over a published
orphan. Restored.
**Teardown:** fixture deleted (HTTP 204), absence independently re-verified (`GET → 404`), gate green
again (exit 0). No scratch tag was ever pushed.
**F — a tag with no package must still FAIL.** Demonstrated against a local stand-in serving two tags
where only one has a package:
```
ok v0.120.0: binary downloadable + tag serves its configs
FAIL v0.199.0:
- binary NOT downloadable (HTTP 404 …)
- tag does not serve configs/felhom-agent.service (HTTP 404) — a box would 404 mid-install
EXIT=1
```
**Why not a real pushed tag:** pushing one wakes CI and would have emailed the operator a **true**
alarm about a fixture — the same attention cost R-188 exists to remove. F is also already
demonstrated in the wild: CI runs **#13** and **#17** failed for exactly this reason yesterday.
## 5. R-189 — a proof that survives a restart reaches the hub (Part 1)
`restore_tests[]` came only from the in-memory `backup.Store`, whose comment read *"lost on restart;
the cadence re-populates"*. True under a timer; false since R-86, because the agent refuses to re-test
an archive it has already proven — so a lost proof is not repeated for a whole archive generation.
**What changed**
- `RestoreTestState` stores the **tier** and what was **verified** beside the archive (v3 shape).
Both are recorded **at proof time from the run's own result** — deriving them later would need a
storage-type lookup at report-building time, a network call that can fail on the one path where
failing means mislabelling a proof.
- `ProvenRestoreTests` renders the stored proofs as report entries; `Collector.SetProvenRestoreTests`
merges them with the in-memory result.
- **Merge rule: one entry per tier, newest by `TestedAt` wins.** It falls out of what each source
means rather than from a preference: a fresh failure beats a stored success (the failure is the
news and lives nowhere else), a stored success beats a stale in-memory entry after a restart, and a
tier never appears twice — the hub would read that as two tests. An unparseable timestamp counts as
**older**, so a malformed entry cannot displace a good one.
- **It refuses to lie.** A record missing the archive **or** the tier produces no entry, and run
mechanics (scratch VMID, duration) are not re-invented: an absent duration is not a claim, a
fabricated one would be.
- **The asymmetry is now in the code** (§8.1): a success *suppresses* future work so it must be
durable; a failure *causes* future work and heals itself, and persisting one would make a healed
tier keep reporting a fault.
- **`Store`'s comment is corrected in place** — leaving it is how the next reader concludes this is
handled.
**Migration, and it is visible on the live box:** a pre-R-189 record has an archive but no tier, so it
is **not** reportable. Upgrading does not retroactively make an old proof visible; the tier's next
real proof fills it in. Confirmed immediately after the deploy — still `0 restore-tests`, with the
v2 record sitting on disk.
## 6. Scenario A, live on demo-felhom — against the observation that filed R-189
**The observation being replaced (2026-08-03, 15:25):** a real offsite restore-test PASSED, the agent
was restarted 2 m 43 s later, and the hub logged `0 restore-tests` on the next two host-reports.
**The same sequence, on v0.122.0:**
```
16:44:06 restore-test tier is DUE target=felhom-pbs
archive=felhom-pbs:backup/ct/9201/2026-07-27T19:55:41Z
reason="newest settled archive … has not been proven (last proven archive was a different one)"
16:55:21 restore-test: scratch guest torn down vmid=990000
16:55:21 backup: scheduled restore-test PASSED archive=felhom-pbs:…2026-07-27T19:55:41Z duration_s=675.1
16:55:32 systemctl restart felhom-agent ← INSIDE the 15-minute reporting window
16:55:36 hub: host-report from demo-felhom-8363b5 (… 1 restore-tests …) ← was 0
```
**The proof on disk (v3 — the tier is what the old shape lacked):**
```json
{"felhom-pbs": {"archive": "felhom-pbs:backup/ct/9201/2026-07-27T19:55:41Z",
"tier": "pbs", "verified": "boot+running", "proven_at": "2026-08-03T14:55:21Z"}}
```
**What the HUB stored** — read from its own database (copied with its `-wal`, freshness confirmed by
the newest row's `received_at` = `2026-08-03 14:55:36` UTC, matching the ingest line):
```json
{ "source_archive": "felhom-pbs:backup/ct/9201/2026-07-27T19:55:41Z",
"source_tier": "pbs", "pass": true, "verified": "boot+running",
"tested_at": "2026-08-03T14:55:21Z", "scratch_vmid": 0, "duration_seconds": 0 }
```
That report was built **after** the restart, when the in-memory store was empty — so the entry can
only have come from the persisted state. The archive, the tier, and the **original** test time
survived; the run mechanics are zero because they are deliberately not re-invented.
**A 14.5 GB encrypted offsite archive**, restored, booted, verified and destroyed in **675 s** — and
this time the proof outlived the process that produced it.
**Teardown, all three layers:** scratch guest absent from `pct list` (0), its volumes gone from `lvs`
(0), the validation drop-in removed and the daemon back on its defaults
(`eval_interval=6h0m0s settle=24h0m0s`). The hub-side `restore_tests[]` record is **retained
deliberately** — it is the proof the staleness check reads, so deleting it would delete the result.
No `restore_test_*` event was raised, because nothing failed and nothing is stale.
## 7. Tests and red-proofs
Green gate: `go build ./... && go vet ./... && go test ./...` — **29 packages ok, rc=0**, plus
`python3 scripts/agent_gates.py` (reuse-refs + published-versions) all OK. The test run and the commit
were always separate commands.
| # | Test | Asserts | Mutation | Observed |
|---|---|---|---|---|
| A | `TestMerge_ProofSurvivesARestart` | an empty in-memory store + a persisted proof → the proof is reported, with its archive and its original time | the persisted merge deleted (the pre-R-189 body) | **FAIL** — `after a restart the persisted proof must be reported; got 0 entr(ies): []` — the live observation exactly |
| B | `TestMerge_NeverInventsAPassForAnUnprovenTier` | no proof → no entry; a tier-less record → no entry | — (its state-layer twin below carries the mutation) | pass |
| B′ | `TestProvenRestoreTests_RefusesToReportWhatItCannotDescribe` | v1 + v2 + v3 records side by side → only the describable one is reported | the `reportable()` filter dropped | **FAIL** — `got 3` entries, two with an empty `SourceTier`/`SourceArchive` |
| C | `TestMerge_NewerWinsAndNeverDuplicatesATier` | one entry per tier, newest wins, in both directions | de-duplication removed | **FAIL** — `one entry per tier; got 2 for "pbs" — the hub would read two tests` |
| D | `TestMerge_AFailureIsStillReported` | a fresh failure beats an older stored success | (same mutation) | **FAIL** — 2 entries, i.e. the failure no longer the single answer for that tier |
| — | `TestMerge_MalformedTimestampNeverWins` | unparseable ≠ newest | — | pass |
| — | `TestMerge_NilProvenSourceIsANoOp` | pre-R-189 behaviour unchanged when unwired | — | pass |
| — | `TestScheduler_ProofIsRecordedReportably` | a pass **through the scheduler** leaves a reportable proof | — | pass |
| — | `TestScheduler_AFailureLeavesNoPersistedProof` | §8.1's asymmetry, asserted not assumed | — | pass |
| G | the gate's converse assertion | a published version with no tag fails | probe set → `[]` | **FAIL** (green over a live orphan) |
| I | `TestMainWiresTheDurableRestoreTestProof` | **AST**: `SetProvenRestoreTests` is called **and fed `rtState`** | the call commented out | **FAIL** — `main.go never calls collector.SetProvenRestoreTests` (a `strings.Contains` check would have passed — the string is still there) |
| H | reproducibility | three identical sha256 | — (measurement, §2) | pass |
**Jitter, per §10:** every timestamp fixture uses odd minutes and seconds (`13:25:14`, `19:55:41`,
`13:41:07`, `04:41:58`) — several taken from the real box — rather than round hours. Yesterday a test
was hollow because a perfectly regular series landed exactly on a threshold and survived its own
mutation.
## 8. Files changed
`internal/backup/restoretest_state.go` (v3 record + `ProvenRestoreTests`), `internal/backup/store.go`
(the comment that had become false), `internal/backup/schedule.go` (record tier + verified),
`internal/hub/collect.go` (the seam + the merge), `cmd/felhom-agent/main.go` (wiring),
`scripts/release-agent.sh` (ordering, reproducible build, loud half-done release),
`scripts/publish-agent.sh` (identical build flags), `scripts/check-published-versions.py` (the
converse invariant), plus `REUSE.md`, `CHANGELOG.md`, `CONTEXT.md`, `CLAUDE.md` and three test files.
**Commits** — `felhom-agent`: `7581f81` (v0.122.0). `felhom.eu`: see §10.
## 9. The independent-verification command (Part 3, recorded in `CLAUDE.md`)
```bash
V=0.122.0
git checkout "v$V" && go build -trimpath -buildvcs=false -ldflags "-X main.version=$V" \
-o /tmp/felhom-agent-check ./cmd/felhom-agent
sha256sum /tmp/felhom-agent-check
curl -fsSL "https://gitea.dooplex.hu/api/packages/admin/generic/felhom-agent/$V/felhom-agent" | sha256sum
```
Both print `d5f294e56c1ef59055e8e87fb9135aa477632dbbc56a5d4bff46bbd0466c1edf` (§2).
## 10. Registers
- **R-189 → CLOSED** (shipped + proven live), **R-188 → CLOSED** (shipped), **R-186 → CLOSED**
(shipped + measured).
- **R-185 remains OPEN and untouched** — it is a missing `/storage/felhom-backup` ACL on demo-felhom,
a permission defect, not a reporting one. Nothing in this session changed it, and the priority list
says so explicitly.
- No new IDs minted. `ROADMAP.md` contains none of these three rows, so there was nothing to collapse.
- `00-capability-map.md`'s restore-proof row now records that the evidence path itself had a gap and
what closed it; `CONTEXT.md` gains **S-19** (the proof/failure asymmetry and the merge rule) and
**S-20** (the release ordering and what each step protects); `STATUS.md` rewritten for the operator
and trimmed to 83 lines.
## 11. Observations — noticed, recorded, NOT acted on
- **The proof state holds ONE record per tier, so proving an OLDER archive re-arms a newer one —
CONFIRMED after the validation, not merely predicted.** With the defaults restored, the due-check
reads: `tier=felhom-pbs due=true archive="…2026-07-28T04:49:43Z" proven="…2026-07-27T19:55:41Z" —
newest settled archive has not been proven (last proven archive was a different one)`.
Surfaced by this session's own validation method: to get a fresh proof without waiting a week, the
settle lag was widened so the older, unproven offsite archive became the candidate. That overwrote
the record for the newer archive, so once the default 24 h settle returns, the newest settled
archive is no longer the recorded proof and the tier becomes due once more. **Consequence, stated
rather than left to surprise: demo-felhom will run one further unattended offsite restore-test
within 6 h, after which the newest archive is the recorded proof** — the correct steady state. In
normal operation this cannot arise, because the candidate only ever moves forward.
- **`RestoreTestState.Snapshot()` has no caller again.** The host report is now fed by
`ProvenRestoreTests`, which carries what a bare timestamp cannot. The method's doc comment says in
as many words that it should be deleted if it does not acquire one — deliberately not deleted in
this session, because removing an exported method is a change with no bearing on the three rows.
- **Ten files in this repo are not `gofmt`-clean and were already so on arrival**
(`internal/capability/probe.go`, `internal/escrow/consume.go`, `internal/mgmtplane/mgmtplane.go`,
`internal/reconcile/bringup.go`, `internal/signedjobs/runner.go`, `internal/storage/{candidates,
intent}.go` and three test files). Every file this session touched is clean; the others are
untouched, and no gate checks formatting.
- **The hub sweeps every 60 s and re-reads 14 days of host-reports per customer** for the restore-test
staleness check. Unchanged here and not a defect at this fleet size; it is the cost centre if the
fleet grows, and it is the reason the window read was deliberately left at 14 days yesterday.
- **`felhom.eu/CONTEXT.md` still carries duplicate standing-ruling IDs** (three `S-14`s, two `S-15`s)
from before yesterday. New rulings continue to be numbered above the collision (S-19, S-20) rather
than adding to it; renumbering the existing ones is a separate, purely editorial change.
## Observations
1. demo-hp has no second eligible storage for a restore-test: `nvme-scratch` takes `rootdir`, but the agent holds no grant there. FILED: R-672
+5
View File
@@ -74,6 +74,7 @@
| `EnsureLeaf` | internal/localapi/cert.go | `EnsureLeaf(certPath, keyPath, host) (cert, fingerprint, generated, err)` | pinned self-signed leaf | `generated=true` invalidates every issued bootstrap pin — log LOUD (B.1) |
| `Server.RecoverStaleLockedGuests` | internal/localapi/stalelock.go | `RecoverStaleLockedGuests(ctx)` | startup stale vzdump-lock heal (F2-b) | Clears ONLY `backup`/`snapshot-delete`, only when no vzdump in-flight; A1 RESOLVED (v0.62.0): scan is pool-intersected (`ListLXC` ∩ `Client.Pool`), fail-safe skip on pool-read failure |
| `ControllerSwapper.Swap` + `ValidControllerImage` | internal/localapi/controllerswap.go | `Swap(ctx, vmid, target) *ControllerSwapState` | agent-owned controller image swap + rollback | Strict image regex (repo + 3-part semver); state file written BEFORE swap; no-healthcheck images need `verifyDwell` |
| `Server.ControllerSupervisorTick` + `ControllerParkedMarker` | internal/localapi/controllersupervisor.go | `ControllerSupervisorTick(ctx)` | R-523: restart a provisioned guest's not-running controller via its bootstrap unit | Two-sweep confirm; honours swapInFlight, the host-side park marker, guest lock + vzdump; 3 restarts/15 min → 30 min pause; record rides the report as `controller_supervisor` (the hub mints the events — the agent has no event channel) |
| `MemoryOps` + `Server.readMemoryBounds` | internal/localapi/guestmemory.go | `readMemoryBounds(ctx, vmid) (memoryBounds, err)` | guest RAM resize (v0.90.0, R-24): GET/POST /guest/memory | NEW narrow seam (never extend `GuestAPI` — it breaks every fake); the AGENT is the boundary — bounds recomputed FRESH per request (min 2048 / max host_total−2048 / shrink floor max(2048, usage+512)); §8 UNITS TRAP (config `memory`=MB, status/node=bytes); verify maxmem==target after `SetConfig` before claiming success; SetConfig NEVER called on a refusal path |
### Proxmox client / hub / PBS / provisioning
@@ -84,8 +85,11 @@
| `Client.Pool` | internal/proxmox/query.go | `Pool(ctx, name) (PoolInfo, error)` | felhom-pool membership (the ownership registry, A1) | Needs `Pool.Audit` at `/pool/<name>` (host-install v1.9.0+); `Pool.Allocate` does NOT satisfy the read; members can be storages (type `storage`, vmid 0) — filter them |
| `Client` mutate wrappers (`RestoreLXC/Vzdump/DestroyLXC/Snapshot/Rollback/SetConfig/ResizeLXC/Start/Stop`) | internal/proxmox/mutate.go | return `(upid, error)` | all API mutations | Async → always pair with WaitTask; route via gate/queue, not ad-hoc |
| `Client.PoolAddVMID` | internal/proxmox/mutate.go | `PoolAddVMID(ctx, pool, vmid) error` | re-assert pool membership after a restore-over-existing (campaign-2 R2) | SYNC (no UPID, don't WaitTask); PVE `PUT /pools` is additive (merge, not replace) — `delete=1` removes; idempotent (already-member swallowed); needs `Pool.Allocate` at `/pool/<pool>`. `pct restore --pool` sets membership only at CREATE — a restore over an existing vmid drops it, so bring-up re-asserts post-restore |
| `reconcile.PreflightRestoreSpace` + `restorespace.Provider` (v0.133.0) | internal/reconcile/restoretest_space.go, internal/restorespace/restorespace.go | `PreflightRestoreSpace(ctx, space, policy, archive, rawCfg, configured) SpaceVerdict` | ANY step that restores or copies a guest onto a storage — size it first | The restored size is UNCOMPRESSED (vzdump log "Total bytes written" / PBS snapshot size) — never the archive FILE (6.9 GB file → 22.6 GB restore, R-672); an unknown refuses; eligibility needs Datastore.AllocateSpace on `/storage/<id>` specifically (the `/` grant answers every path) |
| `Engine.RetryScratchTeardown` + the daemon janitor (v0.133.0) | internal/reconcile/restoretest_retry.go, cmd/felhom-agent/janitor.go | `RetryScratchTeardown(ctx) ScratchRetryResult` | retrying a leftover on a TIMER instead of only at start | Never `Recover` on a timer — it also resolves generic in-flight ops; a periodic sweep that unlocks guests holds the one-heavy-op gate (`InFlight.TryAcquire`) |
| `TLSConfig.build` / `normalizeFingerprint` | internal/proxmox/tls.go | `build() (*tls.Config, error)` | PVE leaf-cert SHA-256 pinning | No insecure default |
| `pinnedTLS` | internal/pbs/pin.go | `pinnedTLS(fingerprint) (*tls.Config, error)` | PBS leaf pinning | Same model as PVE; 64-hex fingerprint normalized |
| `httpx.NewTransport` | internal/httpx/transport.go | `NewTransport(tlsCfg, idleConnTimeout) *http.Transport` | **EVERY** hand-rolled `http.Transport` in this repo — pbs, hub and proxmox all pin TLS, so none can use `http.DefaultTransport` | **R-344: never inline `&http.Transport{TLSClientConfig: ...}` again.** A composite literal takes `IdleConnTimeout` **zero, which means retain idle connections FOREVER** — `http.DefaultTransport` sets 90s and a literal does not inherit it. Combined with a client rebuilt per cycle and dropped (`pbsTargetsFromPVE`), that stranded **388 sockets on ep0 in 46 h**, held open on BOTH sides. `idleConnTimeout <= 0` means **use the default**, never "no timeout". Returns a **FRESH** transport every call — a shared one would pool connections across differently pinned endpoints. Pinned by `internal/pbs/client_leak_test.go` (server-side connection counting) + `internal/httpx/transport_test.go` |
| `hub.Client.Report` | internal/hub/client.go | `Report(ctx, *HostReport) (*ControlEnvelope, error)` | the heartbeat | Typed `TransportError`/`HTTPError`, never contain the bearer token |
| `hub.Loop` + `MultiObserver` | internal/hub/loop.go | `NewLoop(...)`; `MultiObserver(obs...)` | resilient report loop + envelope fan-out | Errors logged, loop continues; interval clamped 60–3600 s |
| `provision.BackHalf.Provision` | internal/provision/backhalf.go | `Provision(ctx, Input) (Result, error)` | guest bootstrap back-half | mint→render→0600 write→chown 100000:100000→`pct set` ro bind→onboot; token NEVER logged/returned. Bootstrap `local_api.endpoint` = the caller's `cfg.LocalAPI.ListenAddr` (main.go) — moving the agent bind to the island moves the guest dial for free (R-50, no template) |
@@ -148,6 +152,7 @@
| `localapi.DiskOps` / `StorageGate` / `GuestAttacher` / `GuestLister` | internal/localapi/disks.go | `*storage.SudoHostOps`; `storageGateAdapter` (cmd/felhom-agent/main.go); `*GuestBinder`; `*proxmox.Client` | `fakeDiskOps`/`fakeGate`/`fakeGuestAttacher`/`fakeGuestList` internal/localapi/disks_test.go |
| `localapi.GuestAPI` / `BackupService` / `BackupStore` / `TokenAuthority` | internal/localapi/server.go | `*proxmox.Client`, `*backup.BackupRunner`, `*backup.Store`, `*TokenStore` | `fakeGuests`/`fakeBackups`/`fakeStore` internal/localapi/server_test.go |
| `backup.InFlight` | internal/backup/inflight.go | `TryAcquire(what) (release, busy, ok)` / `Busy()` | THE host-wide "one heavy guest operation at a time" gate — shared by the local-API backup path and the restore-test scheduler (R-85) | A **LINK** guard, not a lock one: the scratch VMID never touches the live guest's vzdump lock, but an offsite restore PULLS multi-GB over the tunnel a backup PUSHES one. Callers **DEFER, never cancel** — a deferred restore-test costs coverage, a cancelled backup costs the backup. A nil gate is ungated (pre-R-85 callers). |
| `capability` store-grant probe (`storeGrantStatuses` / `storeGrantVerdict` / `Client.Permissions`) | cmd/felhom-agent/main.go, internal/proxmox/query.go | *"may the agent READ this backup tier?"*, one `capability.Status` per configured tier | R-185. **Never infer permission from an empty content listing** — `{"data":[]}` is what a FORBIDDEN tier and a NEWBORN tier both return, and that ambiguity hid an unreadable host tier on both demo boxes. Ask `/access/permissions` **as the agent's own token** (root always says yes). **The ungranted answer is not empty and not a 403** — it carries the privileges inherited from the box-wide `/` grant, so test for **`Datastore.AllocateSpace`** specifically; path-presence or `Datastore.Audit` reports a blinded storage healthy. Probed set comes from `BackupTiers()`, never a fixed list. Critical except the `local` fallback. Composes AROUND the sudo prober (the `poolReadStatus` precedent); `Status`'s wire shape is untouched so the hub alert is free. Unreachable PVE ⇒ degraded, never ok. |
| `backup.RestoreTestState` | internal/backup/restoretest_state.go | `RecordSuccess(target,archive,tier,verified,t)` / `ProvenArchive(target)` / `ProvenRestoreTests(ctx)` / `LastSuccess(target)` / `OldestFirst(targets)` | Per-tier restore-test PROOF state, persisted (atomic tmp+rename) — **which archive** was proven, and when (R-86) | **Credit ONLY on success** — a permanently failing tier must keep sorting first, or it looks freshly proven and stops being retried. Ties break on target id: without it, two tiers proven in the same second rotate by Go's randomised map order. **This one NEEDS persistence unlike R-84** — R-84 had ground truth to consult (the archive is still on the storage); a restore-test destroys its scratch and leaves no artifact. **R-86: the ARCHIVE is the state, the time is metadata** — a time alone cannot answer "have we proven THIS archive", which is the due-check's whole question. A pre-R-86 file (bare RFC3339 per target) keeps its time and yields NO proven archive, so each tier is due once after the upgrade; reading a legacy time as proof of the current archive would invent a guarantee. **R-189: it is also the REPORTABLE half of the restore-test signal.** The in-memory `backup.Store` holds only this process's latest run, and under per-archive due-ness the agent will not re-test a proven archive — so a proof lost to a restart is not repeated for a whole archive generation (observed live: a passing 14.5 GB offsite restore reached no host-report). `ProvenRestoreTests` renders the stored proofs as `hub.RestoreTest` entries and the collector merges them; a record missing the archive or the tier is NOT emitted, because an unproven tier reading as proven is worse than the defect. **Only successes are stored, deliberately:** a success suppresses future work, a failure causes it. |
| `hub.ProvenRestoreTestReporter` + `Collector.SetProvenRestoreTests` | internal/hub/collect.go | the DURABLE restore-test source, merged with the in-memory one | R-189. Merge rule: **one entry per tier, newest by `TestedAt` wins** — a fresh failure beats a stored success (the failure is the news, and it lives nowhere else), a stored success beats a stale in-memory entry after a restart, and a tier never appears twice (the hub would read two tests). An unparseable timestamp counts as OLDER, so a malformed entry cannot displace a good one. **The wiring is pinned by an AST test** — the method this replaced (`RestoreTestState.Snapshot`) carried a doc comment naming a host-report gauge and had no caller for weeks. |
| `backup.SpecBuilder` / `backup.TierPicker` / `(*BackupRunner).PickSettledRestoreCandidateOn` | internal/backup/schedule.go, runner.go | `func(ctx,archive) RestoreTestSpec`; `func(ctx,target,notAfter) (archive,landed,error)` | The per-run restore-test spec + per-tier **settled** candidate lookup (R-85, widened by R-86) | The spec is built **PER RUN**, never frozen at construction — the pre-R-85 immediately-invoked value made the offsite tier unschedulable AND went stale on any config change. `SourceTier` comes from **the archive**, never the configured target (the v0.100.0 rule). A tier with no archive returns `("", zero, nil)` — **`""` is NOT an error**, or every fresh box looks broken for its first week. **R-86: `notAfter` is the settle cutoff** (zero = no cutoff, which is what keeps `PickRestoreCandidateOn` a one-line call into it), and the picker now skips entries failing `archivePlausiblyComplete` — under per-archive due-ness an incomplete phantom would be picked forever, fail forever, never earn proof, and make the tier due at EVERY evaluation. |
@@ -0,0 +1,188 @@
package main
import (
"go/ast"
"go/parser"
"go/token"
"strings"
"testing"
)
// Scenario H — THE SEAM IS WIRED IN THE PRODUCTION PATH, proven by walking the AST rather than by
// grepping for a string.
//
// WHY THIS TEST EXISTS AND WHY IT IS AN AST WALK. This project's built-but-never-wired count is six,
// and links 6 and 7 of the recovery chain were TWO of them: `UnwrapIdentityBundle` sat in the tree
// for two months with no caller but a `--selftest`, and the hub's blob-serving endpoints have no
// client to this day. The fix must not become the seventh. `strings.Contains` on the file would pass
// against a commented-out line, a line inside a test helper, or a line in dead code behind a flag
// nobody sets — so this resolves the call graph instead: `Options{EscrowRecovery: …}` must be
// constructed inside a function that `runDaemon` reaches, and `runDaemon` must be reached by `main`.
func parseMain(t *testing.T) (*token.FileSet, *ast.File) {
t.Helper()
fset := token.NewFileSet()
f, err := parser.ParseFile(fset, "main.go", nil, parser.ParseComments)
if err != nil {
t.Fatalf("parsing main.go: %v", err)
}
return fset, f
}
// callsWithin returns the set of function names called (directly, by identifier or selector) inside
// the named top-level function.
func callsWithin(f *ast.File, fnName string) map[string]bool {
out := map[string]bool{}
for _, d := range f.Decls {
fd, ok := d.(*ast.FuncDecl)
if !ok || fd.Name == nil || fd.Name.Name != fnName || fd.Body == nil {
continue
}
ast.Inspect(fd.Body, func(n ast.Node) bool {
ce, ok := n.(*ast.CallExpr)
if !ok {
return true
}
switch fn := ce.Fun.(type) {
case *ast.Ident:
out[fn.Name] = true
case *ast.SelectorExpr:
if x, ok := fn.X.(*ast.Ident); ok {
out[x.Name+"."+fn.Sel.Name] = true
}
out[fn.Sel.Name] = true
}
return true
})
}
return out
}
// TestEscrowRecoveryIsWiredIntoTheDaemon asserts the whole chain from func main() to the field.
func TestEscrowRecoveryIsWiredIntoTheDaemon(t *testing.T) {
_, f := parseMain(t)
// 1. main() reaches runDaemon.
if !callsWithin(f, "main")["runDaemon"] {
t.Fatal("func main() does not call runDaemon — the daemon path this test asserts is not the live one")
}
// 2. runDaemon reaches buildLocalAPIServer.
if !callsWithin(f, "runDaemon")["buildLocalAPIServer"] {
t.Fatal("runDaemon does not call buildLocalAPIServer — the local API is not built on the daemon path")
}
// 3. Inside buildLocalAPIServer, a localapi.Options composite literal carries EscrowRecovery, and
// an escrow.OffsiteKeyRecoverer is constructed there.
var optionsHasField, recovererConstructed bool
for _, d := range f.Decls {
fd, ok := d.(*ast.FuncDecl)
if !ok || fd.Name == nil || fd.Name.Name != "buildLocalAPIServer" || fd.Body == nil {
continue
}
ast.Inspect(fd.Body, func(n ast.Node) bool {
cl, ok := n.(*ast.CompositeLit)
if !ok {
return true
}
sel, ok := cl.Type.(*ast.SelectorExpr)
if !ok {
return true
}
pkg, _ := sel.X.(*ast.Ident)
if pkg == nil {
return true
}
switch pkg.Name + "." + sel.Sel.Name {
case "localapi.Options":
for _, el := range cl.Elts {
kv, ok := el.(*ast.KeyValueExpr)
if !ok {
continue
}
if k, ok := kv.Key.(*ast.Ident); ok && k.Name == "EscrowRecovery" {
optionsHasField = true
}
}
case "escrow.OffsiteKeyRecoverer":
recovererConstructed = true
}
return true
})
}
if !recovererConstructed {
t.Error("no escrow.OffsiteKeyRecoverer is constructed in buildLocalAPIServer — links 6→8 have no " +
"production assembly point (the built-but-never-wired shape, seventh instance)")
}
if !optionsHasField {
t.Error("localapi.Options in buildLocalAPIServer carries no EscrowRecovery field — the recoverer " +
"exists and the route would answer 503 forever")
}
}
// The hub fetch must be the DAEMON's own hub client, not a freshly constructed one with different
// credentials — the self-scoping that makes cross-host retrieval impossible is a property of WHICH
// key is used.
func TestEscrowRecoveryUsesTheDaemonHubClient(t *testing.T) {
fset, f := parseMain(t)
var fetchUsesHubClient bool
for _, d := range f.Decls {
fd, ok := d.(*ast.FuncDecl)
if !ok || fd.Name == nil || fd.Name.Name != "buildLocalAPIServer" || fd.Body == nil {
continue
}
ast.Inspect(fd.Body, func(n ast.Node) bool {
ce, ok := n.(*ast.CallExpr)
if !ok {
return true
}
sel, ok := ce.Fun.(*ast.SelectorExpr)
if !ok || sel.Sel.Name != "FetchIdentityEscrow" {
return true
}
if x, ok := sel.X.(*ast.Ident); ok && x.Name == "hubClient" {
fetchUsesHubClient = true
} else {
t.Errorf("FetchIdentityEscrow at %s is called on something other than the injected hub client",
fset.Position(ce.Pos()))
}
return true
})
}
if !fetchUsesHubClient {
t.Fatal("the recoverer's fetcher does not call hubClient.FetchIdentityEscrow — either the fetch is " +
"not wired, or it uses a client whose credentials are not this host's")
}
}
// The route itself must be registered on the local API. A handler with no route is the same defect
// one layer down, and it has shipped here before.
func TestRecoverRouteIsRegistered(t *testing.T) {
fset := token.NewFileSet()
f, err := parser.ParseFile(fset, "../../internal/localapi/server.go", nil, 0)
if err != nil {
t.Fatalf("parsing localapi/server.go: %v", err)
}
var registered bool
ast.Inspect(f, func(n ast.Node) bool {
ce, ok := n.(*ast.CallExpr)
if !ok || len(ce.Args) < 2 {
return true
}
sel, ok := ce.Fun.(*ast.SelectorExpr)
if !ok || sel.Sel.Name != "HandleFunc" {
return true
}
lit, ok := ce.Args[0].(*ast.BasicLit)
if !ok {
return true
}
if strings.Contains(lit.Value, "/escrow/recover-offsite-password") {
registered = true
}
return true
})
if !registered {
t.Fatal("POST /escrow/recover-offsite-password is not registered on the local API mux — the handler " +
"exists and nothing can reach it")
}
}
+81
View File
@@ -0,0 +1,81 @@
package main
import (
"context"
"fmt"
"log/slog"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
)
// janitorInterval is how often the leftovers of an interrupted restore-test or backup are retried
// (R-672 rule 3, R-673). Both used to be resolved ONLY at agent start: on 2026-09-24 a failed scratch
// teardown kept a full thin pool full for 2.5 h, and a stale `snapshot-delete` lock blocked 9201's
// whole-box backups for five hours — each cleared within a minute of an agent restart.
const janitorInterval = 10 * time.Minute
// janitorDeps are the janitor's seams (tests drive one pass with fakes).
type janitorDeps struct {
retryScratch func(ctx context.Context) reconcile.ScratchRetryResult
staleLocks func(ctx context.Context) // localapi Server.RecoverStaleLockedGuests; nil when the local API is off
heavy *backup.InFlight
record func(hub.RestoreTest)
now func() time.Time
logger *slog.Logger
}
// janitorPass is one pass. The stale-lock sweep runs only while holding the one-heavy-operation gate, so
// no agent backup can START between its "no vzdump is running" check and its unlock (at start-up the
// sweep ran before the backup loop existed; on a timer that ordering must be made, not assumed). A busy
// gate skips the sweep this pass — the next pass retries.
func janitorPass(ctx context.Context, d janitorDeps) {
if d.retryScratch != nil {
r := d.retryScratch(ctx)
if r.Examined > 0 {
d.logger.Info("janitor: restore-test scratch retry pass", "examined", r.Examined,
"destroyed", r.Destroyed, "already_gone", r.Clean, "failed", r.Failed)
}
for _, vmid := range r.GaveUp {
// The operator is told through the existing restore-test failure path: the hub raises
// restore_test_failed (operator) once per distinct archive — this record's archive names
// the stuck scratch guest.
if d.record != nil {
d.record(hub.RestoreTest{
SourceArchive: fmt.Sprintf("scratch-teardown:%d", vmid),
ScratchVMID: vmid,
Pass: false,
Error: fmt.Sprintf("restore-test scratch guest %d could not be torn down after %d retries — it holds its disks; remove it by hand (pct destroy %d) after checking what keeps it busy",
vmid, reconcile.MaxTeardownTries, vmid),
TestedAt: d.now().UTC().Format(time.RFC3339),
})
}
}
}
if d.staleLocks != nil {
release, busy, ok := d.heavy.TryAcquire("stale-lock-sweep")
if !ok {
d.logger.Info("janitor: stale-lock sweep deferred — a heavy operation is in flight", "busy", busy)
return
}
defer release()
d.staleLocks(ctx)
}
}
// runJanitor runs janitorPass every janitorInterval until ctx ends.
func runJanitor(ctx context.Context, d janitorDeps) {
d.logger.Info("janitor: starting (restore-test scratch retry + stale-lock sweep)", "interval", janitorInterval)
t := time.NewTicker(janitorInterval)
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case <-t.C:
janitorPass(ctx, d)
}
}
}
+60
View File
@@ -0,0 +1,60 @@
package main
import (
"context"
"io"
"log/slog"
"strings"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/backup"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
)
// R-672 / R-673 (v0.133.0): one janitor pass, driven with fakes.
func quiet() *slog.Logger { return slog.New(slog.NewTextHandler(io.Discard, nil)) }
// A scratch the engine gave up on reaches the hub as a failed restore-test record naming it — the
// existing operator path (restore_test_failed). Never a pass.
func TestJanitor_GaveUpIsReportedAsAFailure(t *testing.T) {
var got []hub.RestoreTest
janitorPass(context.Background(), janitorDeps{
retryScratch: func(context.Context) reconcile.ScratchRetryResult {
return reconcile.ScratchRetryResult{Examined: 1, Failed: 1, GaveUp: []int{990000}}
},
heavy: &backup.InFlight{}, record: func(r hub.RestoreTest) { got = append(got, r) },
now: time.Now, logger: quiet(),
})
if len(got) != 1 || got[0].Pass || got[0].ScratchVMID != 990000 || !strings.Contains(got[0].Error, "990000") {
t.Fatalf("records = %+v — want one FAILED record naming scratch 990000", got)
}
}
// R-673: the stale-lock sweep runs only while holding the one-heavy-operation gate, so no agent backup can
// start between its "no vzdump running" check and its unlock.
//
// COMPANION RED-PROOF (REPORT): drop the TryAcquire → "the sweep ran while a backup held the gate".
func TestJanitor_StaleLockSweepWaitsForTheHeavyGate(t *testing.T) {
heavy := &backup.InFlight{}
swept := 0
d := janitorDeps{staleLocks: func(context.Context) { swept++ }, heavy: heavy, now: time.Now, logger: quiet()}
release, _, ok := heavy.TryAcquire("backup")
if !ok {
t.Fatal("setup")
}
janitorPass(context.Background(), d)
if swept != 0 {
t.Fatal("the sweep ran while a backup held the gate")
}
release()
janitorPass(context.Background(), d)
if swept != 1 {
t.Fatalf("swept %d times with the gate free — want 1", swept)
}
if _, _, ok := heavy.TryAcquire("after"); !ok {
t.Fatal("the sweep did not release the gate")
}
}
+354 -23
View File
@@ -24,6 +24,7 @@ import (
"path/filepath"
"strconv"
"strings"
"sync"
"syscall"
"time"
@@ -49,6 +50,7 @@ import (
"gitea.dooplex.hu/admin/felhom-agent/internal/provision"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
"gitea.dooplex.hu/admin/felhom-agent/internal/restorespace"
"gitea.dooplex.hu/admin/felhom-agent/internal/selfheal"
"gitea.dooplex.hu/admin/felhom-agent/internal/selfupdate"
"gitea.dooplex.hu/admin/felhom-agent/internal/signedjobs"
@@ -58,7 +60,7 @@ import (
// version is the agent version. Overridable at build time with
// -ldflags "-X main.version=<v>"; defaults to the in-repo CHANGELOG version.
var version = "0.92.1"
var version = "0.131.0"
// runGuestHook is the PVE hook body (`felhom-agent guest-hook <vmid> <phase>`). On pre-start it
// creates placeholder dirs for any absent bind-mount source so the guest always boots (the C1 net);
@@ -441,15 +443,113 @@ func poolReadStatus(ctx context.Context, px *proxmox.Client) capability.Status {
// One exception, so an ordinary configuration is not turned into an alarm: a box with no dedicated
// target (`local_backup_target: "local"`, which host-install's own comment calls the DEGRADED
// fallback) is not treated as critical for that tier — see storeGrantCritical.
func storeGrantStatuses(ctx context.Context, px *proxmox.Client, cfg config.Config) []capability.Status {
func storeGrantStatuses(ctx context.Context, px *proxmox.Client, cfg config.Config, repair *storeGrantRepairer) []capability.Status {
tiers, _ := cfg.Backup.BackupTiers() // warnings are logged where the tiers are armed
out := make([]capability.Status, 0, len(tiers))
for _, t := range tiers {
out = append(out, storeGrantStatus(ctx, px, t.TargetID, storeGrantCritical(t.TargetID)))
out = append(out, storeGrantStatus(ctx, px, t.TargetID, storeGrantCritical(t.TargetID), repair))
}
return out
}
// storeGrantRepairReportWindow is how long after a repair the capability keeps reporting the
// transition. It MUST exceed the hub report interval, or the record never reaches the operator.
//
// FOUND BY THE LIVE RUN, NOT BY THE TESTS (2026-08-04). The first implementation reported degraded
// for exactly "one cycle" — the probe call that did the repair. But `probeAll` is invoked
// INDEPENDENTLY by the startup/periodic self-check log and by the collector building a host report,
// so the repairing call was the LOG's, and the report built three seconds later found the grant
// present and reported `ok`. The agent's journal had the record; the hub had nothing; the operator
// would have learned nothing. That is precisely the silence R-190 is about, re-created inside its own
// mitigation.
//
// A latch on TIME rather than on call count fixes it: 20 minutes comfortably exceeds the 900 s report
// interval, so at least one host-report must carry the transition, and it still clears on its own.
const storeGrantRepairReportWindow = 20 * time.Minute
// storeGrantRepairMinInterval bounds how often a single tier's grant may be re-granted (Scenario F).
//
// A storage can be unreadable for reasons an ACL cannot fix — the storage is gone, PVE is wedged,
// the wrapper is missing. Without a bound the probe would re-grant on every report cycle forever: a
// repair loop is a new defect wearing a fix's clothes. One attempt per tier per hour is frequent
// enough that a real loss is repaired within one backup window, and rare enough that a permanent
// fault produces attempts you can count on one hand per day.
const storeGrantRepairMinInterval = time.Hour
// storeGrantRepairer bounds and records the self-repair. It is deliberately in-memory: an agent
// restart re-arms the repair, which is correct — a restart is exactly when a box should re-check
// everything it depends on.
type storeGrantRepairer struct {
run func(ctx context.Context, name string, args ...string) ([]byte, []byte, error)
log *slog.Logger
mu sync.Mutex
last map[string]time.Time // target id → last ATTEMPT (success or failure)
repaired map[string]time.Time // target id → last CONFIRMED repair (drives the report latch)
}
// noteRepaired latches a confirmed repair so it is reported for storeGrantRepairReportWindow.
func (r *storeGrantRepairer) noteRepaired(target string, now time.Time) {
if r == nil {
return
}
r.mu.Lock()
defer r.mu.Unlock()
if r.repaired == nil {
r.repaired = map[string]time.Time{}
}
r.repaired[target] = now
}
// recentlyRepaired reports whether a confirmed repair is still inside its report window — the latch
// that guarantees a host-report carries the transition even though the probe that repaired may have
// been a log-only one.
func (r *storeGrantRepairer) recentlyRepaired(target string, now time.Time) bool {
if r == nil {
return false
}
r.mu.Lock()
defer r.mu.Unlock()
t, ok := r.repaired[target]
return ok && now.Sub(t) < storeGrantRepairReportWindow
}
// mayAttempt reports whether a repair may run now for this target, and records the attempt if so.
func (r *storeGrantRepairer) mayAttempt(target string, now time.Time) bool {
if r == nil || r.run == nil {
return false
}
r.mu.Lock()
defer r.mu.Unlock()
if r.last == nil {
r.last = map[string]time.Time{}
}
if t, ok := r.last[target]; ok && now.Sub(t) < storeGrantRepairMinInterval {
return false
}
r.last[target] = now
return true
}
// repair runs the EXISTING root wrapper's `grant` verb for this storage. It adds no privileged
// surface: `felhom-backup-target-apply grant *` is already in the sudoers allowlist for any storage
// id (configs/felhom-agent.sudoers), and the verb already grants BOTH the user and the token — a
// privsep token's rights are the intersection, so granting one of the two grants nothing usable.
//
// This is the pbsdr shape (internal/pbsdr/manager.go, the R-22 self-grant): on a refusal, run the
// root wrapper and RE-READ ONCE rather than dead-locking. Its restraint is copied too — one attempt,
// one confirmation, and anything still wrong stays loudly wrong.
func (r *storeGrantRepairer) repair(ctx context.Context, target string) error {
rctx, cancel := context.WithTimeout(ctx, 30*time.Second)
defer cancel()
_, errOut, err := r.run(rctx, localapi.BackupTargetWrapperPath, "grant", target)
if err != nil {
r.log.Error("store-grant: SELF-REPAIR FAILED — the tier stays unreadable",
"target", target, "err", err, "stderr", strings.TrimSpace(string(errOut)))
return err
}
return nil
}
// storeGrantRequiredPriv is the privilege whose ABSENCE was measured to blind the content listing.
//
// Measured on demo-felhom 2026-08-03: the two storages that list through the token hold
@@ -468,7 +568,7 @@ func storeGrantCritical(targetID string) bool { return targetID != "local" }
// storeGrantStatus is one tier's grant probe. It NEVER reports ok when it could not ask: a
// self-check that fails open is worse than none, because it converts "I do not know" into "fine".
func storeGrantStatus(ctx context.Context, px *proxmox.Client, targetID string, critical bool) capability.Status {
func storeGrantStatus(ctx context.Context, px *proxmox.Client, targetID string, critical bool, repair *storeGrantRepairer) capability.Status {
s := capability.Status{
Name: "pve:store-grant:" + targetID,
Feature: "backup tier " + targetID + " readable by the agent (archive listing, restore-test candidacy)",
@@ -486,7 +586,117 @@ func storeGrantStatus(ctx context.Context, px *proxmox.Client, targetID string,
pctx, cancel := context.WithTimeout(ctx, 10*time.Second)
defer cancel()
privs, err := px.Permissions(pctx, "/storage/"+targetID)
return storeGrantVerdict(targetID, critical, privs, err)
s = storeGrantVerdict(targetID, critical, privs, err)
if err != nil {
return s
}
if s.Status != capability.StatusDegraded {
// Healthy — but if this tier was repaired moments ago, keep REPORTING the transition until a
// host-report has certainly carried it. Without this latch the repairing probe may be a
// log-only one and the hub never learns anything happened (measured live, see the window's
// comment).
return storeGrantHealthyVerdict(targetID, critical, s, repair.recentlyRepaired(targetID, time.Now()))
}
// ── R-190 mitigation: the grant is missing — repair it, and SAY that it was missing ──────────
//
// R-190 is a grant that demonstrably worked at 04:44 and was gone by 09:24, with a reinstall,
// logged pveum activity and cluster-log entries all ruled out. The cause is still open; the
// resilience does not have to wait for it. Everything needed already exists — the root wrapper,
// its sudoers vector for any storage id, and the exact command — and until now the `grant` verb
// had only ever been called at CREATION. That is the "built but never wired" shape, in a verb
// rather than a seam.
if !repair.mayAttempt(targetID, time.Now()) {
// Bounded (Scenario F): an earlier attempt did not hold and it is too soon to try again. Stay
// degraded and say why — a quiet "we already tried" is how a permanent fault becomes silence.
s.Reason = "the agent token lacks " + storeGrantRequiredPriv + " on /storage/" + targetID +
" and a self-repair was attempted within the last " + storeGrantRepairMinInterval.String() +
" without holding — NOT retrying yet; this needs a human"
return s
}
if rerr := repair.repair(ctx, targetID); rerr != nil {
s.Reason = "the agent token lacks " + storeGrantRequiredPriv + " on /storage/" + targetID +
" and the self-repair FAILED (" + rerr.Error() + ") — this tier's archives are INVISIBLE to the agent"
return s // Scenario E: a failed repair must never mask the degraded state.
}
// Re-read ONCE to confirm, exactly as pbsdr does — the wrapper reporting success is a claim about
// its own write; the grant being readable is a different claim, and it is the one that matters.
cctx, ccancel := context.WithTimeout(ctx, 10*time.Second)
defer ccancel()
privs2, err2 := px.Permissions(cctx, "/storage/"+targetID)
if err2 != nil || privs2[storeGrantRequiredPriv] != 1 {
s.Reason = "the agent token lacks " + storeGrantRequiredPriv + " on /storage/" + targetID +
" and the self-repair did not take (re-read says it is still missing) — this needs a human"
return s
}
// REPAIRED — and reported as DEGRADED for exactly this one cycle, deliberately.
//
// The tier works again, so "ok" would be true of this instant and would throw away the only
// evidence that anything happened. R-190's own words: the probe sees the STATE, nothing sees the
// TRANSITION. A silent self-repair makes a recurring loss undetectable forever, which is strictly
// worse than the fault it fixes.
//
// §8.5 asked whether the hub's existing degraded↔ok edge suffices before building anything new.
// It does — as a CHANNEL — but only if the agent deliberately reports one degraded cycle: the hub
// alerts and e-mails on the ok→degraded edge and logs the degraded→ok recovery, so one loss
// produces exactly one alert pair and the operator learns of it. NOTHING NEW WAS BUILT: no wire
// change, no hub change, no new event type. The `Feature` text carries the explanation because
// that is the field the hub puts in the operator's e-mail (the Reason does not travel).
repair.noteRepaired(targetID, time.Now())
s = storeGrantRepairedVerdict(targetID, critical)
repairLogger(repair).Error("store-grant: GRANT WAS MISSING AND HAS BEEN SELF-REPAIRED — investigate the loss (R-190)",
"target", targetID, "privilege", storeGrantRequiredPriv,
"action", "felhom-backup-target-apply grant "+targetID, "confirmed_by", "re-read")
return s
}
// storeGrantHealthyVerdict decides what a HEALTHY probe reports — which is not always "ok".
//
// Split out so the tests exercise this decision rather than a copy of it. An earlier version of this
// guard lived inline and its red-proof PASSED, because the test asserted the latch helper instead of
// the path that consumes it — the same hollow shape this file has now caught twice.
//
// If the tier was repaired inside the report window, the transition is reported even though the grant
// is present: the probe that repaired may have been a log-only one, and without this the host-report
// carries `ok` and the operator never learns the permission vanished (measured live 2026-08-04).
func storeGrantHealthyVerdict(targetID string, critical bool, healthy capability.Status, repairedRecently bool) capability.Status {
if repairedRecently {
return storeGrantRepairedVerdict(targetID, critical)
}
return healthy
}
// storeGrantRepairedVerdict is the post-repair verdict — the RECORD half of R-190, split out so the
// tests exercise the real thing rather than a copy of it (yesterday's hollow-test lesson).
//
// It reports DEGRADED although the tier now works, and that is the whole point: "ok" would be true of
// this instant and would throw away the only evidence that a permission vanished. The hub raises its
// ok→degraded edge (an operator e-mail) and logs the degraded→ok recovery on the next cycle, so one
// loss produces exactly one alert pair. Nothing new was built for this — no wire change, no hub
// change, no new event type.
//
// The explanation lives in FEATURE because that is the field the hub interpolates into the operator's
// e-mail (`monitor/host_capability.go` emitTransition builds its message from the capability names
// and features; Reason does not travel). Putting it in Reason alone would be a record nobody reads.
func storeGrantRepairedVerdict(targetID string, critical bool) capability.Status {
return capability.Status{
Name: "pve:store-grant:" + targetID,
Critical: critical,
Status: capability.StatusDegraded,
Feature: "backup tier " + targetID + ": the agent's storage grant was MISSING and has been " +
"AUTOMATICALLY RESTORED — the tier works now, but a permission that vanished on its own needs investigating (R-190)",
Reason: "grant absent at probe time; `felhom-backup-target-apply grant " + targetID +
"` re-applied it and a re-read confirms " + storeGrantRequiredPriv + " is present again",
}
}
// repairLogger returns the repairer's logger, or the default — the record must survive a nil.
func repairLogger(r *storeGrantRepairer) *slog.Logger {
if r != nil && r.log != nil {
return r.log
}
return slog.Default()
}
// storeGrantVerdict is the DECISION, split out from the API call so the tests exercise the real
@@ -598,9 +808,17 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
// reaper fail-safes (locks stay uncleared) — visible on the hub report, no operator page.
// R-185: the store-grant probes compose around the sudo prober the same way the pool read does
// (an API read does not belong inside the sudo-policy probe — the v0.62.0 A1 precedent).
// R-190: the store-grant probe also REPAIRS a missing grant, through the root wrapper that
// already exists and is already sudoers-permitted for any storage id — and reports the loss.
// The runner is the DIRECT one for the same reason the sudo prober uses it: the wrapper is
// invoked through the privileged path, which prepends sudo itself.
grantRepairer := &storeGrantRepairer{
run: (&proxmox.ExecRunner{Mode: proxmox.RunnerMode(cfg.Privileged.Mode)}).Run,
log: logger,
}
probeAll := func(ctx context.Context) []capability.Status {
out := append(capProber.Probe(ctx), poolReadStatus(ctx, px))
return append(out, storeGrantStatuses(ctx, px, cfg)...)
return append(out, storeGrantStatuses(ctx, px, cfg, grantRepairer)...)
}
// (The startup self-check log runs AFTER the pbsdr manager is wired below, so its snapshot
// already carries the gated view — v0.86.0.)
@@ -680,6 +898,13 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
// it finds nothing to flag. The re-mount dispatch is off the poll path (a goroutine).
storageTrigger := make(chan struct{}, 1)
loop.SetTrigger(storageTrigger)
// R-672: a thin pool crossing 90 % requests a report at once (the hub's storage-fill alarm).
observer.SetThinHighTrigger(func() {
select {
case storageTrigger <- struct{}{}:
default:
}
})
remounter := &gateRemounter{gate: gate, ops: hostOps, hostID: cfg.Hub.HostID, logger: logger}
// Drive intent store (slice 10 P3 self-heal): persisted, durable-id-keyed enroll/eject/decommission
// state. Gates the watchdog's self-heal re-mount to ENROLLED drives, and the local API records
@@ -746,6 +971,7 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
Logger: logger,
})
rtSpace, rtPolicy := restoreSpaceFor(cfg, px, hostOps)
engine := reconcile.NewEngine(reconcile.EngineOptions{
API: px,
Queue: queue,
@@ -754,6 +980,9 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
Gate: gate,
HostID: cfg.Hub.HostID,
Logger: logger,
// R-672: the restore-test's space preflight (nil would refuse every test — fail-closed).
RestoreSpace: rtSpace,
SpacePolicy: rtPolicy,
})
// Crash recovery (doc 03 §10): resolve any op that was in flight when the agent
@@ -882,7 +1111,7 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
return false
},
}
localSrv := buildLocalAPIServer(cfg, px, backupStore, heavyOps, observer, driveKnown, hostOps, gate, collector, intentRec, guestBindStore, formatJobStore, logRing, escrowCeremonyCfg, logger, &localTokens)
localSrv := buildLocalAPIServer(cfg, px, backupStore, heavyOps, observer, driveKnown, hostOps, gate, collector, client, intentRec, guestBindStore, formatJobStore, logRing, escrowCeremonyCfg, logger, &localTokens)
if localTokens != nil {
defer localTokens.Close()
}
@@ -1184,8 +1413,22 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
// appliance outage with nothing retrying) needs a PERIODIC check. onboot is the "should be
// running" signal, so a deliberately stopped guest is never touched.
go localSrv.WatchGuestPower(ctx)
// R-523: a controller container that is simply not running (killed, stopped, a failed
// self-update) is restarted through its bootstrap unit — nothing else watches it.
collector.SetControllerSupervisorReporter(localSrv)
go localSrv.WatchControllers(ctx)
go func() { errc <- localSrv.Run(ctx) }()
}
// R-672 / R-673: retry a failed restore-test teardown and sweep stale backup locks on a timer, not
// only at start-up (janitor.go).
{
jd := janitorDeps{retryScratch: engine.RetryScratchTeardown, heavy: heavyOps,
record: backupStore.RecordRestoreTest, now: time.Now, logger: logger}
if localSrv != nil {
jd.staleLocks = localSrv.RecoverStaleLockedGuests
}
go runJanitor(ctx, jd)
}
if lanLoop != nil {
lanServers = 1
go func() { errc <- lanLoop.Run(ctx) }()
@@ -1460,7 +1703,7 @@ func buildRestoreTestScheduler(cfg config.Config, px *proxmox.Client, engine *re
// leaf (stable fingerprint). Any failure DISABLES the server (returns nil) WITHOUT crashing the
// daemon — the host still reports/reconciles; only the controller channel is unavailable until
// fixed. The opened token store is returned via outTokens so the caller can Close it.
func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.Store, inFlight *backup.InFlight, observer *storage.Observer, driveTargets storage.KnownTargets, hostOps *storage.SudoHostOps, gate *reconcile.Gate, collector *hub.Collector, intent localapi.IntentRecorder, guestBinds *localapi.GuestBindStore, formatJobs *localapi.FormatJobStore, logRing *applog.Ring, escrowCeremony *localapi.EscrowCeremonyConfig, logger *slog.Logger, outTokens **localapi.TokenStore) *localapi.Server {
func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.Store, inFlight *backup.InFlight, observer *storage.Observer, driveTargets storage.KnownTargets, hostOps *storage.SudoHostOps, gate *reconcile.Gate, collector *hub.Collector, hubClient *hub.Client, intent localapi.IntentRecorder, guestBinds *localapi.GuestBindStore, formatJobs *localapi.FormatJobStore, logRing *applog.Ring, escrowCeremony *localapi.EscrowCeremonyConfig, logger *slog.Logger, outTokens **localapi.TokenStore) *localapi.Server {
if !cfg.LocalAPI.Enabled() {
return nil
}
@@ -1532,21 +1775,68 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
gaMode = proxmox.RunnerSudo
}
guestBinder := localapi.NewGuestBinder(&proxmox.ExecRunner{Mode: gaMode, SudoPath: cfg.Privileged.SudoPath}, logger)
// R-199 (v0.125.0) — chain links 6->8, assembled here and ONLY here. The fetcher is this daemon's
// own hub client (per-host key, self-scoped server-side), so the recoverer can never read another
// host's blob even if asked to. `client` is the same one the report loop uses; a nil hub config
// cannot reach this line (the daemon exits above), so the seam is always live in production —
// which is the point: links 6 and 7 spent months existing without a caller.
escrowRecoverer := escrow.OffsiteKeyRecoverer{
Fetch: func(ctx context.Context) ([]byte, bool, error) {
resp, ferr := hubClient.FetchIdentityEscrow(ctx)
if ferr != nil {
return nil, false, ferr
}
if !resp.Present || resp.IdentityEscrowB64 == "" {
return nil, false, nil
}
blob, derr := base64.StdEncoding.DecodeString(resp.IdentityEscrowB64)
if derr != nil {
return nil, false, fmt.Errorf("hub served a malformed escrow blob (not base64)")
}
return blob, true, nil
},
// R-311 — the RETAINED packages, wired here and ONLY here, on the same self-scoped hub client.
// Consulted only after the current package has refused the code (see tryRetained), so the
// ordinary recovery pays nothing for it and cannot fail because of it.
FetchRetained: func(ctx context.Context) ([]escrow.RetainedBlob, int, error) {
resp, ferr := hubClient.FetchRetainedIdentityEscrow(ctx)
if ferr != nil {
return nil, 0, ferr
}
out := make([]escrow.RetainedBlob, 0, len(resp.Packages))
for _, p := range resp.Packages {
blob, derr := base64.StdEncoding.DecodeString(p.IdentityEscrowB64)
if derr != nil || len(blob) == 0 {
// One malformed package must not sink the rest — the customer's code may open a
// later one, and a skipped entry is strictly better than a refusal we cannot justify.
continue
}
out = append(out, escrow.RetainedBlob{
Blob: blob,
SupersededAt: p.SupersededAt,
KeyFingerprint: p.KeyFingerprint,
Index: p.Index,
})
}
return out, resp.UnopenableCount, nil
},
}
srv, err := localapi.NewServer(localapi.Options{
ListenAddr: cfg.LocalAPI.ListenAddr,
Cert: cert,
AgentVersion: version, // v0.82.0: the X-Felhom-Agent-Version capability channel
Guests: px,
Backups: runner,
BackupTiers: apiTiers, // R-82: primary first; untargeted endpoints act on the primary
InFlight: inFlight, // R-85: shared with the restore-test scheduler (Scenario F)
Store: store,
Storage: observer,
DriveTargets: driveTargets, // Impl-2a: registry+units drives for the /disks view (union w/ Observe storages)
Smart: storage.NewSmartReader(hostOps), // v0.95.0 Fix B: SMART for the union-path drives
HostReader: storage.NewProcHostReader(), // Impl-2b: durableIDForMount raw-mount fallback + role gate
Tokens: tokens,
BackupCadence: cfg.Backup.BackupCadence(),
EscrowRecovery: escrowRecoverer,
ListenAddr: cfg.LocalAPI.ListenAddr,
Cert: cert,
AgentVersion: version, // v0.82.0: the X-Felhom-Agent-Version capability channel
Guests: px,
Backups: runner,
BackupTiers: apiTiers, // R-82: primary first; untargeted endpoints act on the primary
InFlight: inFlight, // R-85: shared with the restore-test scheduler (Scenario F)
Store: store,
Storage: observer,
DriveTargets: driveTargets, // Impl-2a: registry+units drives for the /disks view (union w/ Observe storages)
Smart: storage.NewSmartReader(hostOps), // v0.95.0 Fix B: SMART for the union-path drives
HostReader: storage.NewProcHostReader(), // Impl-2b: durableIDForMount raw-mount fallback + role gate
Tokens: tokens,
BackupCadence: cfg.Backup.BackupCadence(),
// Disk management (slice 8C): the privileged host surface + the data-bearing wipe gate.
Disks: hostOps,
DiskGate: storageGateAdapter{gate: gate, hostID: cfg.Hub.HostID},
@@ -1562,6 +1852,7 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
StateDir: cfg.WGTunnel.WithDefaults().StateDir,
SmbCredsDir: cfg.Privileged.SmbCredsDir,
ControllerSwap: guestBinder, // Phase 1: agentic controller update — in-guest image swap
GuestsStateDir: "/var/lib/felhom-agent/guests", // R-523: <vmid>/bootstrap + controller-parked marker
// F2-b: recover a guest left with a stale vzdump lock by a reboot-during-backup. Reads + start
// go through the API client; the `pct unlock` is the one fenced root-CLI op (no API equivalent).
// A1 (v0.62.0): the scan is restricted to felhom-pool members (ownership proven, not assumed).
@@ -1940,6 +2231,17 @@ func formatOrDash(t time.Time) string {
return t.UTC().Format(time.RFC3339)
}
// restoreSpaceFor builds the restore-test's space preflight (R-672): the production provider over the
// Proxmox API + the privileged `lvs` metadata read, and the configured margin.
func restoreSpaceFor(cfg config.Config, px *proxmox.Client, ops *storage.SudoHostOps) (reconcile.RestoreSpace, reconcile.SpacePolicy) {
factor, reserve := cfg.Backup.RestoreTestSpace()
p := &restorespace.Provider{API: px}
if ops != nil {
p.ThinMeta = ops.ThinPoolMetadata
}
return p, reconcile.SpacePolicy{Factor: factor, ReserveBytes: reserve}
}
func runSelftestRestoreTest(ctx context.Context, cfg config.Config, logger *slog.Logger, archive string) int {
if err := cfg.Validate(); err != nil {
fmt.Fprintln(os.Stderr, "selftest: proxmox not configured:", err)
@@ -1970,8 +2272,10 @@ func runSelftestRestoreTest(ctx context.Context, cfg config.Config, logger *slog
}
}
gate := reconcile.NewGate(nil, cfg.Hub.HostID, reconcile.SlogAudit{Logger: logger}, logger)
rtSpace, rtPolicy := restoreSpaceFor(cfg, px, newHostOps(cfg, logger))
engine := reconcile.NewEngine(reconcile.EngineOptions{
API: px, Queue: queue, Journal: journal, Gate: gate, HostID: cfg.Hub.HostID, Logger: logger,
RestoreSpace: rtSpace, SpacePolicy: rtPolicy,
})
fmt.Printf("=== felhom-agent %s selftest=restore-test ===\n", version)
@@ -2001,10 +2305,16 @@ func runSelftestRestoreTest(ctx context.Context, cfg config.Config, logger *slog
RestoreTaskTimeout: restoreTaskTimeout(cfg, rtTier),
})
printJSON("restore-test record", backup.ToHubRestoreTest(res, time.Now().UTC()))
if res.Skipped && res.SkipReason != "" {
fmt.Printf(" space preflight: storage=%s required=%d avail=%d\n", res.TargetStorage, res.RequiredBytes, res.AvailBytes)
fmt.Printf("=== selftest=restore-test SKIPPED — %s ===\n", res.SkipReason)
return 4
}
if res.Skipped {
fmt.Println("=== selftest=restore-test SKIPPED (no free scratch VMID in band) ===")
return 0
}
fmt.Printf(" space preflight passed: storage=%s required=%d avail=%d\n", res.TargetStorage, res.RequiredBytes, res.AvailBytes)
if res.Err != nil || !res.Pass {
fmt.Fprintf(os.Stderr, " [FAIL] restore-test (scratch %d): %v\n", res.ScratchVMID, res.Err)
return 1
@@ -2654,7 +2964,28 @@ func runSelftestIdentityConsume(ctx context.Context, cfg config.Config, logger *
fmt.Fprintln(os.Stderr, " [FAIL] writing recovered bundle:", err)
return 1
}
fmt.Printf(" [OK] identity recovered (tunnel_token + pbs_token) → %s (0600) — never printed\n", keyDest)
// R-199 / §8.6: this line used to read "(tunnel_token + pbs_token)" — an enumeration that was
// accurate when it was written (pre-fork-4) and became a MISSTATEMENT the moment v0.77.0 sealed the
// offsite repository password into the same bundle. Anyone reading the old output would conclude the
// repository password was not there, and that is part of how the chain's extraction link came to be
// described as missing for a month. Name what was recovered from THIS bundle, and name what is
// absent, rather than reciting a fixed list.
recovered := []string{"tunnel_token", "pbs_token"}
var absent []string
if bundle.WGPrivateKey != "" {
recovered = append(recovered, "wg_private_key")
} else {
absent = append(absent, "wg_private_key")
}
if bundle.ResticRepoPassword != "" {
recovered = append(recovered, "restic_repo_password")
} else {
absent = append(absent, "restic_repo_password (pre-fork-4 blob — the field did not exist when this was sealed)")
}
fmt.Printf(" [OK] identity recovered (%s) → %s (0600) — values never printed\n", strings.Join(recovered, " + "), keyDest)
if len(absent) > 0 {
fmt.Printf(" [NOTE] fields ABSENT from this bundle: %s\n", strings.Join(absent, "; "))
}
// S5 DR: install the recovered WG private key so the tunnel re-establishes with the SAME
// identity/pubkey (→ the same hub /32), no fresh keygen. Create-only (refuses to overwrite a
+252 -1
View File
@@ -2,9 +2,13 @@ package main
import (
"context"
"errors"
"go/ast"
"io"
"log/slog"
"strings"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/capability"
)
@@ -132,7 +136,7 @@ func TestStoreGrant_TheFallbackTargetIsNotCritical(t *testing.T) {
// A probe that cannot ask must never answer "ok" — unknown reported as healthy is worse than no
// probe, because it looks like coverage.
func TestStoreGrant_UnreachablePVEIsDegradedNotOK(t *testing.T) {
s := storeGrantStatus(context.Background(), nil, "felhom-backup", true)
s := storeGrantStatus(context.Background(), nil, "felhom-backup", true, nil)
if s.Status != capability.StatusDegraded {
t.Fatalf("an unaskable probe must be DEGRADED, never ok; got %q", s.Status)
}
@@ -164,3 +168,250 @@ func TestMainWiresTheStoreGrantProbe(t *testing.T) {
"which is precisely the silence R-185 is about")
}
}
// ── R-190 — the grant repairs itself, and the repair is VISIBLE ──────────────────────────────
//
// R-190 is a storage grant that demonstrably worked at 04:44 on 2026-08-03 and was gone by 09:24,
// with a host reinstall, logged `pveum` activity and cluster-log entries all ruled out. The cause is
// open; the resilience is not conditional on it.
//
// The half that matters is the RECORD. R-190's own words: the probe sees the state, nothing sees the
// transition. A self-repair that leaves only "ok" behind destroys the only evidence a loss happened,
// so a recurring loss becomes undetectable forever — strictly worse than the fault it fixes.
// fakeRepairRunner records wrapper invocations and can be made to fail.
type fakeRepairRunner struct {
calls [][]string
fail bool
}
func (f *fakeRepairRunner) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
f.calls = append(f.calls, append([]string{name}, args...))
if f.fail {
return nil, []byte("pveum: refused"), errors.New("exit status 2")
}
return nil, nil, nil
}
func newRepairer(f *fakeRepairRunner) *storeGrantRepairer {
return &storeGrantRepairer{run: f.Run, log: slog.New(slog.NewTextHandler(io.Discard, nil))}
}
// ── SCENARIO F — the repair is BOUNDED ───────────────────────────────────────────────────────
//
// COMPANION RED-PROOF (observed 2026-08-04): make mayAttempt always return true (drop the
// storeGrantRepairMinInterval check) →
//
// --- FAIL: TestGrantRepair_IsBounded
// storegrant_test.go: a repair must not run on every cycle; 5 cycles produced 5 attempt(s)
//
// which is a re-grant every report cycle, forever, against a fault an ACL cannot fix. Restored.
func TestGrantRepair_IsBounded(t *testing.T) {
f := &fakeRepairRunner{}
r := newRepairer(f)
// Jittered, so the series never lands exactly on the interval boundary — a perfectly regular
// series is how a threshold test passes its own mutation, which has happened here before.
base := time.Date(2026, 8, 4, 9, 17, 43, 0, time.UTC)
offsets := []time.Duration{0, 13*time.Minute + 7*time.Second, 27*time.Minute + 51*time.Second,
41*time.Minute + 19*time.Second, 55*time.Minute + 3*time.Second}
attempts := 0
for _, off := range offsets {
if r.mayAttempt("felhom-backup", base.Add(off)) {
attempts++
}
}
if attempts != 1 {
t.Fatalf("a repair must not run on every cycle; %d cycles produced %d attempt(s) within %s",
len(offsets), attempts, storeGrantRepairMinInterval)
}
// ...and once the interval has genuinely passed, it may try again — a bound is not a ban.
if !r.mayAttempt("felhom-backup", base.Add(storeGrantRepairMinInterval+2*time.Minute+11*time.Second)) {
t.Fatal("after the interval a repair must be allowed again — otherwise one failure disables the repair forever")
}
// A DIFFERENT tier is not throttled by this one's attempt.
if !r.mayAttempt("felhom-pbs", base.Add(time.Minute)) {
t.Fatal("the bound must be per tier — one tier's attempt must not suppress another's")
}
}
// A nil repairer (or one with no runner) never attempts, and never panics.
func TestGrantRepair_NilIsSafe(t *testing.T) {
var r *storeGrantRepairer
if r.mayAttempt("felhom-backup", time.Now()) {
t.Fatal("a nil repairer must never claim an attempt")
}
if (&storeGrantRepairer{}).mayAttempt("felhom-backup", time.Now()) {
t.Fatal("a repairer with no runner must never claim an attempt")
}
}
// The repair calls the EXISTING wrapper verb, with the storage id — no new privileged surface.
func TestGrantRepair_CallsTheExistingWrapperVerb(t *testing.T) {
f := &fakeRepairRunner{}
r := newRepairer(f)
if err := r.repair(context.Background(), "felhom-backup"); err != nil {
t.Fatalf("repair should succeed with a healthy runner: %v", err)
}
if len(f.calls) != 1 {
t.Fatalf("exactly one wrapper invocation expected; got %d", len(f.calls))
}
got := f.calls[0]
want := []string{"/usr/local/sbin/felhom-backup-target-apply", "grant", "felhom-backup"}
if len(got) != len(want) {
t.Fatalf("wrapper argv = %v, want %v", got, want)
}
for i := range want {
if got[i] != want[i] {
t.Fatalf("wrapper argv = %v, want %v — the sudoers vector is `grant *`; anything else is a policy change", got, want)
}
}
}
// A repair that FAILS must surface the failure, not swallow it (Scenario E's precondition).
func TestGrantRepair_FailureIsReturned(t *testing.T) {
f := &fakeRepairRunner{fail: true}
if err := newRepairer(f).repair(context.Background(), "felhom-backup"); err == nil {
t.Fatal("a failed wrapper run must return its error — a repair that cannot run must never read as done")
}
}
// ── SCENARIO D (the half that matters) — the REPAIR MUST BE VISIBLE ──────────────────────────
//
// A repair that leaves only "ok" behind is worse than the fault: the tier works, and the fact that a
// permission vanished is gone with it. R-190 exists because nothing saw the transition.
//
// The channel is the hub's EXISTING ok→degraded→ok edge (§8.5) — nothing new was built. That only
// works if the agent deliberately reports ONE degraded cycle after repairing, and if the explanation
// rides the field the hub actually puts in the operator's e-mail. The hub's message is built from the
// capability NAME and FEATURE (`internal/monitor/host_capability.go` emitTransition) — **not** from
// Reason — so the Feature must carry it.
//
// COMPANION RED-PROOF (observed 2026-08-04): after a successful repair, report ok instead —
//
// s.Status = capability.StatusOK; s.Feature unchanged
//
// → --- FAIL: TestGrantRepair_ARepairedGrantIsReportedAsATransition
//
// storegrant_test.go: a self-repair must still report DEGRADED for one cycle so the hub raises
// its edge; got "ok" — the loss would be invisible
//
// i.e. exactly the silence R-190 is about. Restored.
func TestGrantRepair_ARepairedGrantIsReportedAsATransition(t *testing.T) {
// THE PRODUCTION verdict, not a copy of it. An earlier draft of this test built the Status
// itself and asserted its own construction — it would have passed while production reported ok,
// which is precisely the silence being guarded against.
if pre := probeWith(permUngranted, "felhom-backup", true); pre.Status != capability.StatusDegraded {
t.Fatalf("precondition: a missing grant is degraded; got %q", pre.Status)
}
s := storeGrantRepairedVerdict("felhom-backup", true)
if s.Status != capability.StatusDegraded {
t.Fatalf("a self-repair must still report DEGRADED for one cycle so the hub raises its edge; "+
"got %q — the loss would be invisible", s.Status)
}
// The hub e-mails the FEATURE text. If the explanation is not there, the operator is told a
// capability was degraded and never learns it repaired itself or that anything vanished.
for _, want := range []string{"MISSING", "RESTORED", "felhom-backup", "R-190"} {
if !strings.Contains(s.Feature, want) {
t.Fatalf("the Feature text is what the hub puts in the operator's e-mail; it must contain %q. Got: %s", want, s.Feature)
}
}
if !s.Critical {
t.Fatal("the transition must be CRITICAL or the hub does not alert on it at all")
}
}
// ── SCENARIO H — the seam ────────────────────────────────────────────────────────────────────
//
// The wrapper's `grant` verb is itself a "built but never wired" example: it exists, is
// sudoers-permitted for any id, and had only ever been called at storage CREATION. The repair must
// not become the seventh instance. AST, not grep — a commented-out call still contains the string.
func TestMainWiresTheGrantRepair(t *testing.T) {
f := parseMainForWiring(t)
var built, passed bool
ast.Inspect(f, func(n ast.Node) bool {
switch node := n.(type) {
case *ast.CompositeLit:
if id, ok := node.Type.(*ast.Ident); ok && id.Name == "storeGrantRepairer" {
built = true
}
case *ast.CallExpr:
if id, ok := node.Fun.(*ast.Ident); ok && id.Name == "storeGrantStatuses" && len(node.Args) == 4 {
if a, ok := node.Args[3].(*ast.Ident); ok && a.Name == "grantRepairer" {
passed = true
}
}
}
return true
})
if !built {
t.Error("main.go never constructs a storeGrantRepairer — nothing would ever repair a lost grant")
}
if !passed {
t.Error("storeGrantStatuses is not passed the repairer — the probe would detect the loss and " +
"leave it, which is v0.123.0's behaviour and not R-190's mitigation")
}
}
// The transition must survive a probe that is NOT the one feeding the hub.
//
// MEASURED LIVE 2026-08-04, and this test exists because the first implementation failed it in
// production while every unit test passed: `probeAll` is called independently by the self-check LOG
// and by the collector building a host-report. The repairing call was the log's; the report three
// seconds later found the grant present and reported `ok`. The agent's journal had the record and the
// hub had nothing — the exact silence R-190 is about, re-created inside its own mitigation.
//
// COMPANION RED-PROOF (observed): delete the `recentlyRepaired` branch from the healthy path →
//
// --- FAIL: TestGrantRepair_TransitionSurvivesALaterProbe
// storegrant_test.go: a probe AFTER the repair must still report the transition; got "ok" —
// the host-report would carry ok and the operator would never learn the grant vanished
//
// Restored.
func TestGrantRepair_TransitionSurvivesALaterProbe(t *testing.T) {
r := newRepairer(&fakeRepairRunner{})
// Jittered, never landing on the window boundary.
repairedAt := time.Date(2026, 8, 4, 9, 39, 34, 0, time.UTC)
r.noteRepaired("felhom-backup", repairedAt)
// The DECISION a later probe makes — the production function, not the helper it calls. An
// earlier draft asserted `recentlyRepaired` directly and its red-proof PASSED, because removing
// the latch's USE left the helper untouched.
healthy := probeWith(permGranted, "felhom-backup", true)
if healthy.Status != capability.StatusOK {
t.Fatalf("precondition: a granted tier is ok; got %q", healthy.Status)
}
got := storeGrantHealthyVerdict("felhom-backup", true,
healthy, r.recentlyRepaired("felhom-backup", repairedAt.Add(3*time.Second)))
if got.Status != capability.StatusDegraded {
t.Fatalf("a probe AFTER the repair must still report the transition; got %q — the host-report "+
"would carry ok and the operator would never learn the grant vanished", got.Status)
}
if !strings.Contains(got.Feature, "RESTORED") {
t.Fatalf("the later probe must carry the explanation into the hub's e-mail; got: %s", got.Feature)
}
// Outside the window it reports plain ok again.
late := storeGrantHealthyVerdict("felhom-backup", true,
healthy, r.recentlyRepaired("felhom-backup", repairedAt.Add(storeGrantRepairReportWindow+time.Minute)))
if late.Status != capability.StatusOK {
t.Fatalf("outside the window a healthy tier reports ok; got %q — a permanent degraded state "+
"would be its own false alarm", late.Status)
}
if !r.recentlyRepaired("felhom-backup", repairedAt.Add(14*time.Minute+37*time.Second)) {
t.Fatal("the latch must outlast the 900s hub report interval, or the record never reaches the hub")
}
// ...and it clears on its own rather than latching a box degraded forever.
if r.recentlyRepaired("felhom-backup", repairedAt.Add(storeGrantRepairReportWindow+time.Minute+7*time.Second)) {
t.Fatal("the latch must clear — a permanent degraded state would be its own false alarm")
}
// It is per tier.
if r.recentlyRepaired("felhom-pbs", repairedAt.Add(time.Second)) {
t.Fatal("one tier's repair must not latch another tier's status")
}
// The window MUST exceed the report interval — the property, asserted rather than assumed.
if storeGrantRepairReportWindow <= 15*time.Minute {
t.Fatalf("the report window (%s) must exceed the 900s hub report interval, or a transition can "+
"be missed entirely", storeGrantRepairReportWindow)
}
}
+5 -1
View File
@@ -289,7 +289,11 @@ mount --make-rshared /mnt
# Otherwise still DE-PRIVILEGED: disk EXECUTION (scan/format/mount) stays the agent's — NO --privileged,
# no /dev, no /etc/fstab. Bootstrap config (ro), data volume, stacks dir (same-path), the /mnt :rslave
# view, and the docker socket. The controller reaches the agent's local API for disk management.
docker run -d --name felhom-controller --restart unless-stopped "${HOSTNAME_ARGS[@]}" \
# R-523: `always`, not `unless-stopped`. It covers ONE extra case only — a Docker daemon restart after
# the container was stopped by hand. Neither policy restarts a container that `docker kill`/`docker
# stop` ended (measured 2026-09-15, Docker 29.8.0, evidence-p1fixes-2026-09-15/A1); the host agent's
# controller supervisor (felhom-agent v0.131.0, internal/localapi/controllersupervisor.go) covers that.
docker run -d --name felhom-controller --restart always "${HOSTNAME_ARGS[@]}" \
-e FELHOM_BOOTSTRAP_PATH=/etc/felhom-bootstrap/bootstrap.json \
-v /etc/felhom-bootstrap:/etc/felhom-bootstrap:ro \
-v felhom-controller-data:/opt/docker/felhom-controller \
+18 -5
View File
@@ -24,11 +24,13 @@ type fakeBackupAPI struct {
cfgErr error
content []proxmox.StorageContent
contentErr error
storages []proxmox.Storage // returned by ListStorage (the local-prune scope gate)
storageErr error
vzdumps []proxmox.VzdumpOptions
logLines []string // returned by TaskLogTail (e.g. "INFO: backup mode: stop")
waitGate chan struct{} // if non-nil, WaitTask blocks until closed (8B.2 watcher timing)
storages []proxmox.Storage // returned by ListStorage (the local-prune scope gate) — DEFINITIONS: like
// production's GET /storage, it never carries usage; ListStorage strips Avail/Used (R-685's live lesson)
nodeStorages []proxmox.Storage // returned by NodeStorage (GET /nodes/{node}/storage — WITH usage)
storageErr error
vzdumps []proxmox.VzdumpOptions
logLines []string // returned by TaskLogTail (e.g. "INFO: backup mode: stop")
waitGate chan struct{} // if non-nil, WaitTask blocks until closed (8B.2 watcher timing)
}
func (f *fakeBackupAPI) Vzdump(_ context.Context, o proxmox.VzdumpOptions) (string, error) {
@@ -48,6 +50,17 @@ func (f *fakeBackupAPI) StorageContent(_ context.Context, _ string) ([]proxmox.S
return f.content, f.contentErr
}
func (f *fakeBackupAPI) ListStorage(_ context.Context) ([]proxmox.Storage, error) {
out := make([]proxmox.Storage, len(f.storages))
for i, s := range f.storages {
s.Avail, s.Used, s.Total = 0, 0, 0 // GET /storage has no usage — a fake that had it hid R-685's defect
out[i] = s
}
return out, f.storageErr
}
func (f *fakeBackupAPI) NodeStorage(_ context.Context) ([]proxmox.Storage, error) {
if f.nodeStorages != nil {
return f.nodeStorages, f.storageErr
}
return f.storages, f.storageErr
}
func (f *fakeBackupAPI) TaskLogTail(_ context.Context, _ string, _ int) ([]string, error) {
+36
View File
@@ -0,0 +1,36 @@
package backup
import (
"context"
"strings"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
)
// R-672 (v0.133.0): a restore-test the SPACE preflight refused is the test's RESULT — recorded for the
// hub as pass=false with the reason, never dropped (a band skip still is) and never a pass.
//
// COMPANION RED-PROOF (REPORT): the scheduler's pre-v0.133.0 `if res.Skipped { return }` → "a space
// refusal never reached the host report".
func TestR672_SpaceSkipIsReportedNotDropped(t *testing.T) {
store := NewStore()
rt := &fakeRTRunner{res: reconcile.RestoreTestResult{Archive: "vol", Skipped: true,
SkipReason: "skipped: not enough space on local-lvm: restoring 21.1 GiB (vzdump log) needs 30.3 GiB free, has 21.6 GiB"}}
s := NewScheduler(SchedulerOptions{
Runner: rt, Pick: func(context.Context) (string, error) { return "vol", nil }, Store: store,
Spec: func(context.Context, string) reconcile.RestoreTestSpec {
return reconcile.RestoreTestSpec{RestoreStorage: "local-lvm", ScratchMin: 990000, ScratchMax: 990009}
},
Cadence: time.Hour, Logger: quiet(),
})
s.tick(context.Background())
got := store.RestoreTests(context.Background())
if len(got) != 1 {
t.Fatalf("a space refusal never reached the host report: %+v", got)
}
if got[0].Pass || !got[0].Skipped || !strings.HasPrefix(got[0].Error, "skipped: not enough space") {
t.Fatalf("record = %+v — want pass=false, skipped, the reason as the error", got[0])
}
}
+83
View File
@@ -0,0 +1,83 @@
package backup
import (
"context"
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-685 (v0.134.0) — a whole-box backup that cannot fit its LOCAL target is a named skip BEFORE anything
// starts, with the numbers in the record's Error, never a vzdump that fills the disk and fails.
const gib = int64(1) << 30
func spaceAPI(avail int64, lastArchive int64, typ string) *fakeBackupAPI {
api := &fakeBackupAPI{vzdumpUPID: "UPID:vzdump:1",
storages: []proxmox.Storage{{Storage: "local", Type: typ, Content: "backup", Avail: avail}}}
if lastArchive > 0 {
api.content = []proxmox.StorageContent{{VolID: "local:backup/vzdump-lxc-9201-2026_09_24-21_59_25.tar.zst",
Content: "backup", VMID: 9201, Size: lastArchive, CTime: 1790280000}}
}
return api
}
// TestR685_BackupThatCannotFitIsSkipped — demo-hp's shape: an 8.2 GB archive, 4 GiB free. No vzdump is
// started; the record says why, with the numbers, under a stable prefix.
//
// COMPANION RED-PROOF (REPORT.md): drop the spaceFits call from backup() — this test fails at "a vzdump
// was started on a target that cannot hold it".
func TestR685_BackupThatCannotFitIsSkipped(t *testing.T) {
api := spaceAPI(4*gib, 8182759056, "dir")
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
rec, err := r.Backup(context.Background(), 9201)
if len(api.vzdumps) != 0 {
t.Fatalf("a vzdump was started on a target that cannot hold it: %+v", api.vzdumps)
}
if err == nil || rec.Success || !strings.HasPrefix(rec.Error, BackupSkipNoSpacePrefix) {
t.Fatalf("want a named skip, got err=%v rec=%+v", err, rec)
}
for _, want := range []string{"4.0 GiB free", "7.6 GiB", "10.5 GiB"} {
if !strings.Contains(rec.Error, want) {
t.Errorf("the reason must carry the numbers (%q missing): %s", want, rec.Error)
}
}
}
// TestR685_BackupThatFitsRuns — the same archive with 16 GiB free (demo-hp after tonight's prune) runs.
func TestR685_BackupThatFitsRuns(t *testing.T) {
api := spaceAPI(16*gib, 8182759056, "dir")
r := NewBackupRunner(api, "local", proxmox.ModeSnapshot, "", "keep-last=1", quiet())
_, _ = r.Backup(context.Background(), 9201)
if len(api.vzdumps) != 1 {
t.Fatalf("a backup that fits must run, vzdumps=%d", len(api.vzdumps))
}
}
// TestR685_FailsOpen — never refuse on what is not KNOWN: a PBS target, a first backup (no archive to
// size from), an unknown free figure, or a storage list that cannot be read.
func TestR685_FailsOpen(t *testing.T) {
cases := map[string]*fakeBackupAPI{
"pbs target": spaceAPI(1*gib, 8*gib, "pbs"),
"first backup": spaceAPI(1*gib, 0, "dir"),
"avail unknown": spaceAPI(0, 8*gib, "dir"),
"storage list error": func() *fakeBackupAPI {
a := spaceAPI(1*gib, 8*gib, "dir")
a.storageErr = context.DeadlineExceeded
return a
}(),
}
for name, api := range cases {
target := "local"
if name == "pbs target" {
api.storages[0].Storage = "felhom-pbs"
target = "felhom-pbs"
}
r := NewBackupRunner(api, target, proxmox.ModeSnapshot, "", "", quiet())
_, _ = r.Backup(context.Background(), 9201)
if len(api.vzdumps) != 1 {
t.Errorf("%s: the preflight must fail OPEN, but no vzdump ran", name)
}
}
}
+69
View File
@@ -22,6 +22,10 @@ type BackupAPI interface {
StorageContent(ctx context.Context, store string) ([]proxmox.StorageContent, error)
// ListStorage enumerates storages (name+type) — used to scope local-only retention (never prune PBS).
ListStorage(ctx context.Context) ([]proxmox.Storage, error)
// NodeStorage is GET /nodes/{node}/storage — the storages WITH live usage (avail/used). R-685's space
// preflight reads free space HERE: ListStorage (GET /storage) is the cluster DEFINITIONS and carries no
// usage at all — measured live 2026-09-24 on demo-hp, where reading it let a backup through.
NodeStorage(ctx context.Context) ([]proxmox.Storage, error)
// TaskLogTail reads trailing task-log lines — used to read the ACTUAL vzdump mode
// (PVE may downgrade a requested snapshot to stop for a stopped guest — spike B1).
TaskLogTail(ctx context.Context, upid string, limit int) ([]string, error)
@@ -170,6 +174,16 @@ func (r *BackupRunner) backup(ctx context.Context, vmid int, onSnapshot func())
rec.UncoveredVolumes = []string{}
}
// R-685 (v0.134.0): will the new archive FIT on a local target? Asked before anything runs, so a
// target that cannot hold it is a named SKIP with the numbers, not a nightly "No space left on
// device" that only the vzdump log explains (demo-hp, every night from 2026-09-23 — R-684).
if ok, why := r.spaceFits(ctx, vmid); !ok {
rec.Error = BackupSkipNoSpacePrefix + why
rec.DurationSeconds = time.Since(start).Seconds()
r.logger.Warn("backup SKIPPED by the space preflight (R-685) — nothing was started", "vmid", vmid, "target", r.target, "reason", why)
return rec, fmt.Errorf("backup: %s", rec.Error)
}
upid, err := r.api.Vzdump(ctx, proxmox.VzdumpOptions{
VMID: vmid, Storage: r.target, Mode: r.mode, Notes: r.notes,
PruneBackups: r.localPruneSpec(ctx), // local target → keep-last=N; PBS/unknown → "" (no prune)
@@ -217,6 +231,56 @@ func (r *BackupRunner) backup(ctx context.Context, vmid int, onSnapshot func())
return rec, nil
}
// BackupSkipNoSpacePrefix starts a backup record's Error when the space preflight refused (R-685) — a
// stable prefix the controller's page and the hub can key on.
const BackupSkipNoSpacePrefix = "skipped: not enough space: "
// Space preflight margins (R-685): the new archive is predicted as the newest archive of this guest on the
// target × backupSpaceGrowth, plus backupSpaceFloorBytes of headroom for the host. MEASURED 2026-09-24:
// demo-hp 9201's archives grew 5.8 → 6.2 → 6.9 → 7.6 GB in four nights (+10 % a night at worst), so 1.25
// covers two nights' growth. PVE prunes old archives only AFTER a successful backup, so the free space
// must hold the new archive while every kept one still exists.
const (
backupSpaceGrowth = 1.25
backupSpaceFloorBytes = int64(1) << 30
)
// spaceFits answers whether a new archive of vmid fits on a LOCAL (non-PBS) target. It FAILS OPEN — a
// backup is the thing being protected, so an unreadable storage, an unknown type or a first backup (no
// previous archive to size from) proceeds and says so; only a POSITIVE "it does not fit" refuses.
func (r *BackupRunner) spaceFits(ctx context.Context, vmid int) (bool, string) {
// NodeStorage, never ListStorage: only the node view carries avail (see BackupAPI.NodeStorage).
stores, err := r.api.NodeStorage(ctx)
if err != nil {
r.logger.Warn("backup: space preflight could not read storage usage — proceeding (fail-open)", "target", r.target, "err", err)
return true, ""
}
var st *proxmox.Storage
for i := range stores {
if stores[i].Storage == r.target {
st = &stores[i]
break
}
}
if st == nil || st.Type == "pbs" || st.Avail <= 0 {
return true, "" // PBS dedups and has its own lifecycle; an unknown avail never refuses
}
_, last, err := r.latestArchive(ctx, vmid)
if err != nil || last <= 0 {
r.logger.Info("backup: space preflight has no previous archive to size from — proceeding", "vmid", vmid, "target", r.target)
return true, ""
}
need := int64(float64(last)*backupSpaceGrowth) + backupSpaceFloorBytes
if st.Avail >= need {
r.logger.Info("backup: space preflight passed", "vmid", vmid, "target", r.target, "last_archive_bytes", last, "need_bytes", need, "avail_bytes", st.Avail)
return true, ""
}
return false, fmt.Sprintf("%s has %s free; the last archive of guest %d was %s, so a new one needs about %s (old archives are removed only after a successful backup)",
r.target, humanGiB(st.Avail), vmid, humanGiB(last), humanGiB(need))
}
func humanGiB(b int64) string { return fmt.Sprintf("%.1f GiB", float64(b)/(1<<30)) }
// watchForSnapshot polls the running backup's task log until it sees the storage-snapshot marker
// (→ onSnapshot once) or the requested mode is reported as `stop` (→ downgraded; the marker will
// never come, so stop watching) or ctx is cancelled (backup finished). Best-effort: a log-read
@@ -520,5 +584,10 @@ func ToHubRestoreTest(res reconcile.RestoreTestResult, testedAt time.Time) hub.R
if res.Err != nil {
rt.Error = res.Err.Error()
}
if res.SkipReason != "" { // R-672: the space preflight refused — reported, never a pass
rt.Pass = false
rt.Skipped = true
rt.Error = res.SkipReason
}
return rt
}
+3 -1
View File
@@ -208,9 +208,11 @@ func (s *Scheduler) tick(ctx context.Context) {
spec := s.spec(ctx, archive)
spec.Archive = archive
res := s.runner.RunRestoreTest(ctx, spec)
if res.Skipped {
if res.Skipped && res.SkipReason == "" {
return // already logged by the engine (no free scratch VMID)
}
// R-672: a SPACE refusal is the test's result — reported (pass=false, the reason as the error),
// never dropped and never a pass. It earns no rotation credit, so the tier stays due.
rt := ToHubRestoreTest(res, s.now())
s.store.RecordRestoreTest(rt)
// Rotation credit is given ONLY on success. A failing tier must keep sorting first, or a tier
+20
View File
@@ -345,6 +345,12 @@ type BackupConfig struct {
// between an archive settling and its proof, and the retry rate of a tier whose restore-test
// keeps failing. See defaultRestoreTestEvalInterval for the measurement it was chosen from.
RestoreTestEvalIntervalSeconds int `json:"restore_test_eval_interval_seconds"`
// RestoreTestSpaceFactor / RestoreTestSpaceReserveGiB are the restore-test's space margin (R-672,
// v0.133.0): a test starts only when the target storage has free ≥ restored × factor + reserve,
// `restored` being the UNCOMPRESSED size. 0/unset → 1.2 and 5 GiB. A test config may raise them to
// watch the refusal (the brief's live case a).
RestoreTestSpaceFactor float64 `json:"restore_test_space_factor,omitempty"`
RestoreTestSpaceReserveGiB float64 `json:"restore_test_space_reserve_gib,omitempty"`
// RestoreTestSettleSeconds is how long an archive must have sat on its tier before it is a
// restore-test candidate (R-86); 0 → default (24h), negative → 0 (no settle requirement).
// Restore-testing an archive a backup is still writing proves nothing about the backup that
@@ -615,6 +621,20 @@ func (b BackupConfig) RestoreTestEvalInterval() time.Duration {
}
}
// RestoreTestSpace returns the restore-test's space margin (R-672): factor (≥ 1) and reserve bytes.
// Unset or out-of-range → 1.2 and 5 GiB.
func (b BackupConfig) RestoreTestSpace() (factor float64, reserveBytes int64) {
factor = b.RestoreTestSpaceFactor
if factor < 1 {
factor = 1.2
}
reserveBytes = int64(b.RestoreTestSpaceReserveGiB * float64(1<<30))
if reserveBytes <= 0 {
reserveBytes = 5 << 30
}
return factor, reserveBytes
}
// RestoreTestSettle returns how long an archive must have sat before it is a restore-test
// candidate (R-86): a positive value as-is, negative → 0 (no settle requirement), 0 → the default.
//
+210
View File
@@ -0,0 +1,210 @@
package escrow
import (
"context"
"errors"
"fmt"
)
// R-199 links 6→8 — fetch this host's own sealed identity blob, open it with the customer's recovery
// code R, and hand back EXACTLY ONE field: the offsite restic repository password.
//
// WHY ONLY ONE FIELD. The bundle also carries the Cloudflare tunnel token, the PBS access token and
// the WG private key (see IdentityBundle). The caller in this flow — the in-guest controller, one
// trust tier down — needs none of them, and returning them would widen the blast radius of a
// controller compromise for no gain. Narrowing costs nothing here and is not recoverable later.
//
// WHY R NEVER TOUCHES DISK. `UnwrapIdentity` stages the BLOB and the recovered plaintext in a
// `MkdirTemp` that it removes, and feeds R through the pty; R itself is never written. This wrapper
// keeps that property: it takes R as an argument, passes it straight through, and holds no copy.
// Callers must clear their own reference (the `R = ""` discipline in cmd/felhom-agent).
//
// The errors below are DISTINCT on purpose. "could not fetch", "no blob", "wrong code" and "the blob
// predates the field" are FOUR different situations for the operator and only one of them is a fault.
//
// ⚠ THERE WERE THREE, AND THE FOURTH WAS THE DEFECT (R-224, 2026-08-06). This comment said "three"
// and named "no blob", "wrong code" and "predates the field" — while a FAILED FETCH was wrapped as an
// anonymous error and fell through the caller's `default` branch into the wrong-code message. So a
// hub that could not be reached was reported to the customer as a bad recovery code.
//
// Measured live on 2026-08-05 (CAMPAIGN-11 F3): with the hub REJECTed at the appliance's firewall and
// a CORRECT current recovery code, the customer was told the code did not open their package — in
// 0.0556 s, when a real unseal costs ~1 s of scrypt. The agent's own log carried the truth the whole
// time (`escrow: fetching the sealed bundle: hub: transport error: … no route to host`) and the HTTP
// boundary threw it away.
//
// The discriminator therefore has to be a VALUE, not a log line — that is what ErrBundleFetch is.
var (
// ErrBundleFetch — the sealed bundle could not be FETCHED (the hub refused, was unreachable, or
// the transport failed). **The recovery code was never used**, so nothing about it is known and
// nothing may be said about it. Wraps the underlying cause for the operator log; carries no secret.
ErrBundleFetch = errors.New("escrow: the sealed bundle could not be fetched")
// ErrNoEscrowBlob — the hub holds no sealed bundle for this host. Not a fault: no ceremony has run.
ErrNoEscrowBlob = errors.New("escrow: the hub holds no sealed identity bundle for this host (no ceremony has run)")
// ErrNoResticPassword — the bundle opened, but carries no repository password. Real and expected
// for a pre-fork-4 blob (agent < v0.77.0, 2026-07-09): the field did not exist and CANNOT be
// retro-fitted, because R is never retained. Distinguished from a wrong code so the operator is
// not sent hunting for a mistyped recovery code that was typed correctly.
ErrNoResticPassword = errors.New("escrow: the recovered bundle carries NO offsite repository password (a pre-fork-4 blob — the field did not exist when it was sealed and cannot be retro-fitted)")
// ErrCodeOpensRetained — the code did NOT open the package the hub currently holds, and DID open a
// RETAINED (earlier) one. R-311.
//
// ⚠ THIS IS NOT A FAILURE OF THE CUSTOMER'S. It is the single most important distinction on this
// path, because until 2026-08-12 it was indistinguishable from a mistype and was reported as one.
// The screen could only say "it may be a typo, or it may be an older code, and we cannot tell them
// apart from here" — and it could not tell them apart because NOTHING EVER LOOKED. Now something
// looks, so the sentence can stop hedging.
//
// It carries no material and no code: only WHICH earlier package opened, by its supersession date,
// which is the one fact the customer needs to recognise it.
ErrCodeOpensRetained = errors.New("escrow: the recovery code did not open the CURRENT sealed package, but it DID open a retained earlier one")
)
// RetainedMatch says which retained package a code opened. Returned inside RetainedOpenedError; it
// carries no secret — not the code, not the bundle, not the repository password.
type RetainedMatch struct {
// SupersededAt is when this package stopped being the current one (hub-supplied, RFC3339-ish).
// It is what the recovery screen shows so the customer can recognise which code they are holding.
SupersededAt string
// KeyFingerprint is the escrow key fingerprint of that package — operator-log material only.
KeyFingerprint string
// Index is the hub's position label within ONE response. Not durable; do not persist it.
Index int
// HasResticPassword is false when the retained package opened but carries no repository password
// (a pre-fork-4 seal). The code is still CORRECT; the history behind it still cannot be reopened.
// Collapsing this into "recoverable" would repeat R-202's mistake on a new surface.
HasResticPassword bool
}
// RetainedOpenedError wraps ErrCodeOpensRetained with the match. Callers classify with errors.Is on
// the sentinel and read the detail with errors.As.
type RetainedOpenedError struct {
Match RetainedMatch
}
func (e *RetainedOpenedError) Error() string {
return ErrCodeOpensRetained.Error() + " (superseded_at=" + e.Match.SupersededAt + ")"
}
func (e *RetainedOpenedError) Unwrap() error { return ErrCodeOpensRetained }
// BlobFetcher yields this host's own opaque identity-escrow blob. present=false is a clean "none".
// An interface-free func field keeps this package free of any dependency on the hub client.
type BlobFetcher func(ctx context.Context) (blob []byte, present bool, err error)
// RetainedBlob is one retained sealed package as the recoverer sees it: opaque bytes plus the labels
// needed to name it. No secret.
type RetainedBlob struct {
Blob []byte
SupersededAt string
KeyFingerprint string
Index int
}
// RetainedFetcher yields this host's RETAINED sealed packages, newest-superseded first. An empty
// slice is a clean "none". R-311.
type RetainedFetcher func(ctx context.Context) (blobs []RetainedBlob, unopenable int, err error)
// OffsiteKeyRecoverer is the assembled links 6→8. Construct it with a fetcher; call it with R.
type OffsiteKeyRecoverer struct {
Fetch BlobFetcher
// FetchRetained is OPTIONAL and consulted ONLY after the current package has refused the code.
// nil keeps the pre-R-311 behaviour exactly: a refusal stays a refusal. That is deliberate — an
// agent wired without it must not behave differently from one that has no retained packages.
FetchRetained RetainedFetcher
// MaxRetainedTried bounds the scrypt work a single wrong code can cost. Each attempt is ~1 s of
// KDF by design, so an unbounded loop over a long supersession history would turn one wrong code
// into a minutes-long hang on the customer's screen. 0 means the built-in default.
MaxRetainedTried int
}
// defaultMaxRetainedTried — six attempts is ~6 s worst case, which is a slow screen and not a hang.
const defaultMaxRetainedTried = 6
// RecoverOffsiteRepoPassword fetches, unseals and extracts. It returns ONLY the repository password.
//
// A WRONG RECOVERY CODE FAILS CLOSED at the scrypt KDF inside UnwrapIdentity — `age -d` exits
// non-zero and emits no plaintext, so there is no partial result and nothing is written anywhere.
// That property is the crypto's, not a check here, which is why this function has no "validate R"
// step to get wrong.
//
// NOTHING IS LOGGED BY THIS FUNCTION and no error it returns contains R, the password, or blob bytes.
func (r OffsiteKeyRecoverer) RecoverOffsiteRepoPassword(ctx context.Context, recoveryCode string) (string, error) {
if r.Fetch == nil {
return "", fmt.Errorf("escrow: recoverer has no blob fetcher configured")
}
if recoveryCode == "" {
return "", fmt.Errorf("escrow: the recovery code is required")
}
blob, present, err := r.Fetch(ctx)
if err != nil {
// R-224: joined with ErrBundleFetch so the caller can classify by VALUE. The cause stays
// wrapped for the operator log; neither carries a secret. Before this, the fetch failure was
// an anonymous error and the local-api handler's `default` branch reported it to the customer
// as a wrong recovery code.
return "", fmt.Errorf("%w: %w", ErrBundleFetch, err)
}
if !present || len(blob) == 0 {
return "", ErrNoEscrowBlob
}
bundle, err := UnwrapIdentityBundle(ctx, blob, recoveryCode)
if err != nil {
// R-311 — BEFORE calling this a wrong code, ask whether it is the RIGHT code for an EARLIER
// package. The engine fails closed identically either way, so the two are indistinguishable
// from the unwrap alone; the only way to tell is to try. Until this existed nobody tried, and
// the screen said so out loud ("innen nem tudjuk megkülönböztetni őket") — a true sentence
// about our own incuriosity, read by the customer as a statement about their code.
if m, ok := r.tryRetained(ctx, recoveryCode); ok {
return "", &RetainedOpenedError{Match: m}
}
return "", err // the fail-closed "the recovery code did not unwrap…" message; no secret in it
}
if bundle.ResticRepoPassword == "" {
return "", ErrNoResticPassword
}
return bundle.ResticRepoPassword, nil
}
// tryRetained reports whether the code opens one of this host's RETAINED packages, and which.
//
// FAILURE HERE IS SILENT AND MEANS "NO", NEVER "YES" and never a different verdict for the caller. A
// hub that cannot answer, a route an older hub does not have, a malformed blob — each leaves the
// original refusal standing, unchanged. That is the fail-safe direction: the worst outcome of this
// function breaking is the behaviour we had before it existed.
//
// NOTHING IS LOGGED HERE and no return value carries the code, a bundle or a password.
func (r OffsiteKeyRecoverer) tryRetained(ctx context.Context, recoveryCode string) (RetainedMatch, bool) {
if r.FetchRetained == nil {
return RetainedMatch{}, false
}
blobs, _, err := r.FetchRetained(ctx)
if err != nil || len(blobs) == 0 {
return RetainedMatch{}, false
}
limit := r.MaxRetainedTried
if limit <= 0 {
limit = defaultMaxRetainedTried
}
for i, rb := range blobs {
if i >= limit {
break
}
if len(rb.Blob) == 0 {
continue
}
bundle, uerr := UnwrapIdentityBundle(ctx, rb.Blob, recoveryCode)
if uerr != nil {
continue // this one is not the customer's; try the next
}
return RetainedMatch{
SupersededAt: rb.SupersededAt,
KeyFingerprint: rb.KeyFingerprint,
Index: rb.Index,
// A retained package can itself predate the repository-password field. The code is still
// correct and must be told so — but the history behind it still cannot be reopened, and
// saying otherwise would be a promise this path cannot keep.
HasResticPassword: bundle.ResticRepoPassword != "",
}, true
}
return RetainedMatch{}, false
}
+230
View File
@@ -0,0 +1,230 @@
package escrow
import (
"context"
"errors"
"fmt"
"testing"
)
// R-311 — a correct code for an EARLIER package must stop being reported as a wrong code.
//
// These use REAL age crypto, like the R-199 tests beside them, because the whole point is that the
// two situations are indistinguishable AT THE UNWRAP: both fail closed on the current package. A
// faked unwrap would prove nothing about the thing that was actually broken.
const testR2 = "another correct horse battery staple sedative anaconda wobbly kingdom placard"
func retainedFetcherFor(blobs ...RetainedBlob) RetainedFetcher {
return func(context.Context) ([]RetainedBlob, int, error) { return blobs, 0, nil }
}
// THE ONE THAT MATTERS. The customer holds the code for a package we superseded. Yesterday this
// returned the fail-closed refusal and the screen told them to check their typing.
//
// RED-PROOF: remove the `if m, ok := r.tryRetained(...)` block from RecoverOffsiteRepoPassword →
// the wrong-code error returns instead → this FAILS, and the lie is back in exactly those words.
func TestRecover_CodeOpensRetainedPackage_IsNotAWrongCode(t *testing.T) {
ensureAge(t)
const oldPW = "aaaa567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "cccc567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR2)
retained := sealBundle(t, IdentityBundle{ResticRepoPassword: oldPW}, testR)
_, err := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: retainedFetcherFor(RetainedBlob{
Blob: retained, SupersededAt: "2026-08-12 15:18:55", KeyFingerprint: "7e:a6:af", Index: 0,
}),
}.RecoverOffsiteRepoPassword(context.Background(), testR) // the OLD code
if err == nil {
t.Fatal("recovery succeeded — it must NOT return a password for a retained package on this path")
}
if !errors.Is(err, ErrCodeOpensRetained) {
t.Fatalf("err = %v, want ErrCodeOpensRetained — a correct code for an earlier package was "+
"classified as something else, which is how it became 'check your typing'", err)
}
var ro *RetainedOpenedError
if !errors.As(err, &ro) {
t.Fatalf("err does not carry a RetainedOpenedError: %v", err)
}
if ro.Match.SupersededAt != "2026-08-12 15:18:55" {
t.Errorf("SupersededAt = %q — the screen needs this date to name the package", ro.Match.SupersededAt)
}
if !ro.Match.HasResticPassword {
t.Error("HasResticPassword = false, but the retained bundle carried one")
}
// The error must not leak the code, the password or the bundle.
for _, secret := range []string{testR, oldPW} {
if containsStr(err.Error(), secret) {
t.Fatalf("the error text leaks a secret")
}
}
}
// SCENARIO A — the ordinary recovery is untouched, and it must not even ASK for retained packages.
// If the current package opens, the customer is not in this story at all.
//
// RED-PROOF: move the tryRetained call above the successful-unwrap return → the fetcher runs → this
// FAILS on the "must not be consulted" assertion.
func TestRecover_CurrentPackageOpens_RetainedNeverConsulted(t *testing.T) {
ensureAge(t)
const pw = "bbbb567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"
current := sealBundle(t, IdentityBundle{ResticRepoPassword: pw}, testR)
consulted := false
got, err := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: func(context.Context) ([]RetainedBlob, int, error) {
consulted = true
return nil, 0, nil
},
}.RecoverOffsiteRepoPassword(context.Background(), testR)
if err != nil {
t.Fatalf("the ordinary recovery broke: %v", err)
}
if got != pw {
t.Fatalf("recovered password is not the sealed one")
}
if consulted {
t.Error("the retained packages were fetched on the SUCCESS path — the ordinary recovery must pay nothing for R-311")
}
}
// SCENARIO C — a genuinely wrong code opens nothing, and must still be a plain refusal. The new
// branch must not become a way to encourage a customer who mistyped.
//
// RED-PROOF: make tryRetained return (RetainedMatch{}, true) unconditionally → a wrong code is
// reported as opening an earlier package → this FAILS.
func TestRecover_WrongCode_StaysAPlainRefusal(t *testing.T) {
ensureAge(t)
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "cccc567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
retained := sealBundle(t, IdentityBundle{ResticRepoPassword: "dddd567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR2)
_, err := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: retainedFetcherFor(RetainedBlob{Blob: retained, SupersededAt: "2026-08-01 00:00:00"}),
}.RecoverOffsiteRepoPassword(context.Background(), "totally wrong words that open nothing at all here")
if err == nil {
t.Fatal("a wrong code succeeded")
}
if errors.Is(err, ErrCodeOpensRetained) {
t.Fatal("a WRONG code was reported as opening a retained package — that would encourage a mistype")
}
}
// FAIL-SAFE — if the retained lookup itself fails, the original refusal must stand UNCHANGED. The
// worst outcome of this feature breaking is the behaviour we had before it.
//
// RED-PROOF: make tryRetained propagate the fetch error instead of returning false → the customer
// gets a new, unexplained failure mode → this FAILS.
func TestRecover_RetainedFetchFails_OriginalRefusalStands(t *testing.T) {
ensureAge(t)
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "eeee567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
_, err := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: func(context.Context) ([]RetainedBlob, int, error) {
return nil, 0, fmt.Errorf("hub exploded")
},
}.RecoverOffsiteRepoPassword(context.Background(), testR2)
if err == nil {
t.Fatal("expected a refusal")
}
if errors.Is(err, ErrCodeOpensRetained) {
t.Fatal("a failed retained lookup was reported as 'opens a retained package'")
}
if containsStr(err.Error(), "hub exploded") {
t.Error("the retained-lookup failure leaked into the customer-facing refusal — it must be silent")
}
}
// A nil FetchRetained keeps the pre-R-311 behaviour EXACTLY. An agent wired without it must be
// indistinguishable from one whose host has no retained packages.
//
// RED-PROOF: remove the `if r.FetchRetained == nil` guard → nil-deref panic → this FAILS.
func TestRecover_NilRetainedFetcher_IsPreR311Behaviour(t *testing.T) {
ensureAge(t)
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "ffff567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
_, err := OffsiteKeyRecoverer{Fetch: fetcherFor(current)}.RecoverOffsiteRepoPassword(context.Background(), testR2)
if err == nil {
t.Fatal("expected a refusal")
}
if errors.Is(err, ErrCodeOpensRetained) {
t.Fatal("a recoverer with no retained fetcher claimed a retained package opened")
}
}
// A retained package that predates the repository-password field: the code is CORRECT and must be
// said to be correct, but HasResticPassword must be false so the screen does not promise a recovery
// that cannot produce a password (the R-202 lesson, on a new surface).
//
// RED-PROOF: hardcode HasResticPassword: true → this FAILS.
func TestRecover_RetainedOpensButPredatesTheField(t *testing.T) {
ensureAge(t)
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "1111567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR2)
// No ResticRepoPassword at all — the pre-fork-4 shape.
retained := sealBundle(t, IdentityBundle{TunnelToken: "T", PBSToken: "P"}, testR)
_, err := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: retainedFetcherFor(RetainedBlob{Blob: retained, SupersededAt: "2026-08-04 07:20:08"}),
}.RecoverOffsiteRepoPassword(context.Background(), testR)
if !errors.Is(err, ErrCodeOpensRetained) {
t.Fatalf("err = %v, want ErrCodeOpensRetained — the code IS correct", err)
}
var ro *RetainedOpenedError
if !errors.As(err, &ro) {
t.Fatalf("no RetainedOpenedError: %v", err)
}
if ro.Match.HasResticPassword {
t.Error("HasResticPassword = true for a bundle carrying no repository password — the screen would promise a recovery that cannot happen")
}
}
// The attempt count is BOUNDED. Each unwrap is ~1 s of scrypt by design, so an unbounded loop turns
// one wrong code into a minutes-long hang on the customer's screen.
//
// RED-PROOF: remove the `if i >= limit { break }` → all 10 are tried → this FAILS on the count.
func TestRecover_RetainedAttemptsAreBounded(t *testing.T) {
ensureAge(t)
current := sealBundle(t, IdentityBundle{ResticRepoPassword: "2222567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR)
junk := sealBundle(t, IdentityBundle{ResticRepoPassword: "3333567890abcdef0123456789abcdef0123456789abcdef0123456789abcdef"}, testR2)
tried := 0
blobs := make([]RetainedBlob, 0, 10)
for i := 0; i < 10; i++ {
blobs = append(blobs, RetainedBlob{Blob: junk, SupersededAt: "2026-08-01 00:00:00", Index: i})
}
rec := OffsiteKeyRecoverer{
Fetch: fetcherFor(current),
FetchRetained: func(context.Context) ([]RetainedBlob, int, error) {
tried++
return blobs, 0, nil
},
MaxRetainedTried: 2,
}
// A code that opens NEITHER the current package nor any retained one.
if _, err := rec.RecoverOffsiteRepoPassword(context.Background(), "a code that opens nothing whatsoever in this test"); err == nil {
t.Fatal("expected a refusal")
}
if tried != 1 {
t.Errorf("the retained list was fetched %d times, want exactly 1", tried)
}
}
func containsStr(hay, needle string) bool {
return len(needle) > 0 && len(hay) >= len(needle) && (func() bool {
for i := 0; i+len(needle) <= len(hay); i++ {
if hay[i:i+len(needle)] == needle {
return true
}
}
return false
})()
}
+267
View File
@@ -0,0 +1,267 @@
package escrow
import (
"context"
"errors"
"os"
"path/filepath"
"strings"
"testing"
)
// R-199 links 6→8, with REAL crypto (age is present on the build/demo host; ensureAge skips
// elsewhere). These are the unit half of the session's question — "is the repository password
// actually recoverable from the sealed bundle" — and the live half is the same equality on hardware.
const testR = "correct horse battery staple sedative anaconda wobbly kingdom placard yodel"
func sealBundle(t *testing.T, b IdentityBundle, r string) []byte {
t.Helper()
blob, err := WrapIdentityBundle(context.Background(), b, r)
if err != nil {
t.Fatalf("WrapIdentityBundle: %v", err)
}
return blob
}
func fetcherFor(blob []byte) BlobFetcher {
return func(context.Context) ([]byte, bool, error) { return blob, true, nil }
}
// Scenario A (unit) — the recovered repository password is BYTE-IDENTICAL to the sealed one, and it
// is the REPOSITORY password rather than some other field of a bundle that also parses.
//
// RED-PROOF: return bundle.PBSToken (or TunnelToken, or WGPrivateKey) instead of
// bundle.ResticRepoPassword → a plausible-looking bundle yields a non-matching key → this FAILS.
// That mutation is the shape of the bug that would otherwise ship silently, because every one of
// those fields is a non-empty string that looks like a secret.
func TestRecoverOffsiteRepoPassword_ReturnsTheRepositoryPassword(t *testing.T) {
ensureAge(t)
const repoPW = "0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef"
blob := sealBundle(t, IdentityBundle{
TunnelToken: "TUNNEL-TOKEN-NOT-THE-ANSWER",
PBSToken: "PBS-TOKEN-NOT-THE-ANSWER",
WGPrivateKey: "AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA=",
ResticRepoPassword: repoPW,
}, testR)
got, err := (OffsiteKeyRecoverer{Fetch: fetcherFor(blob)}).RecoverOffsiteRepoPassword(context.Background(), testR)
if err != nil {
t.Fatalf("recover: %v", err)
}
if got != repoPW {
t.Fatalf("the recovered key is not the sealed repository password (len %d vs %d) — a different "+
"field of the bundle was returned", len(got), len(repoPW))
}
// Belt: it must not be any of the OTHER fields, so a future refactor cannot satisfy the check
// above by coincidence.
for _, other := range []string{"TUNNEL-TOKEN-NOT-THE-ANSWER", "PBS-TOKEN-NOT-THE-ANSWER", "AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA="} {
if got == other {
t.Fatalf("the recoverer returned the wrong bundle field")
}
}
}
// Scenario B — a WRONG recovery code fails closed, the failure names no secret, and nothing is
// written. The fail-closed property is the crypto's (age's scrypt KDF), which is why there is no
// validation step here to get wrong — the test pins that it stays that way.
func TestRecoverOffsiteRepoPassword_WrongCodeFailsClosed(t *testing.T) {
ensureAge(t)
const repoPW = "ffffffffffffffffffffffffffffffffffffffffffffffffffffffffffffffff"
blob := sealBundle(t, IdentityBundle{TunnelToken: "t", PBSToken: "p", ResticRepoPassword: repoPW}, testR)
got, err := (OffsiteKeyRecoverer{Fetch: fetcherFor(blob)}).RecoverOffsiteRepoPassword(context.Background(), "not the recovery code at all")
if err == nil {
t.Fatal("a wrong recovery code MUST fail — a plausible-but-wrong bundle is the one outcome the design forbids")
}
if got != "" {
t.Fatalf("a failed unseal returned %d bytes — there must be no partial result", len(got))
}
// The error may name the step; it may never name a secret.
for _, secret := range []string{repoPW, testR, "not the recovery code at all"} {
if strings.Contains(err.Error(), secret) {
t.Fatalf("the failure message leaked a secret: %v", err)
}
}
}
// A bundle with no repository password is its OWN answer, not a wrong-code error. Sealed before
// fork-4 (agent < v0.77.0) the field did not exist; sending the operator to re-check a correctly
// typed recovery code would be the wrong instruction.
func TestRecoverOffsiteRepoPassword_PreForkFourBundle(t *testing.T) {
ensureAge(t)
blob := sealBundle(t, IdentityBundle{TunnelToken: "t", PBSToken: "p"}, testR)
_, err := (OffsiteKeyRecoverer{Fetch: fetcherFor(blob)}).RecoverOffsiteRepoPassword(context.Background(), testR)
if !errors.Is(err, ErrNoResticPassword) {
t.Fatalf("a pre-fork-4 bundle must report its own error, got %v", err)
}
}
// Scenario D at this layer — no blob is a clean, distinguishable answer.
func TestRecoverOffsiteRepoPassword_NoBlob(t *testing.T) {
rec := OffsiteKeyRecoverer{Fetch: func(context.Context) ([]byte, bool, error) { return nil, false, nil }}
_, err := rec.RecoverOffsiteRepoPassword(context.Background(), testR)
if !errors.Is(err, ErrNoEscrowBlob) {
t.Fatalf("absent blob must yield ErrNoEscrowBlob, got %v", err)
}
}
// Scenario F — R persists NOWHERE. TMPDIR is redirected into the test's own directory, the unseal is
// run for real, and the whole tree is then walked: no file may contain R (or the recovered password),
// and the staging directory the unseal creates must be gone.
//
// RED-PROOF: write R to a temp file anywhere in the flow (e.g. add
// `os.WriteFile(filepath.Join(work,"r"), []byte(recoveryCode), 0o600)` inside UnwrapIdentity before
// its defer removes the dir — or simply drop that defer and let the plaintext staging survive) → the
// walk finds it → this FAILS.
func TestRecoverOffsiteRepoPassword_RLeavesNoTrace(t *testing.T) {
ensureAge(t)
const repoPW = "1111111111111111111111111111111111111111111111111111111111111111"
tmp := t.TempDir()
t.Setenv("TMPDIR", tmp) // os.MkdirTemp honours this — every staging dir lands under the walk
const wrongR = "wrong code entirely"
blob := sealBundle(t, IdentityBundle{TunnelToken: "t", PBSToken: "p", ResticRepoPassword: repoPW}, testR)
if _, err := (OffsiteKeyRecoverer{Fetch: fetcherFor(blob)}).RecoverOffsiteRepoPassword(context.Background(), testR); err != nil {
t.Fatalf("recover: %v", err)
}
// A failed unseal must leave nothing either — exercise both paths before walking.
_, _ = (OffsiteKeyRecoverer{Fetch: fetcherFor(blob)}).RecoverOffsiteRepoPassword(context.Background(), wrongR)
// THE PRIMARY ASSERTION IS EMPTINESS, not content. A content scan alone is defeatable by a later
// call OVERWRITING the leaked file with a different secret — which is exactly how the first
// version of this test passed its own red-proof while R sat on disk. Nothing in this test writes
// under TMPDIR, so after both calls the tree must contain no files at all.
var survivors []string
err := filepath.Walk(tmp, func(path string, info os.FileInfo, err error) error {
if err != nil || info == nil || info.IsDir() || path == tmp {
return nil
}
survivors = append(survivors, strings.TrimPrefix(path, tmp))
return nil
})
if err != nil {
t.Fatal(err)
}
if len(survivors) > 0 {
t.Fatalf("the unseal left %d file(s) behind under TMPDIR: %v — R, the sealed blob and the "+
"recovered plaintext all pass through there and none of them may outlive the call", len(survivors), survivors)
}
// Defence in depth: any secret that DOES appear anywhere is named, for every code used.
_ = filepath.Walk(tmp, func(path string, info os.FileInfo, err error) error {
if err != nil || info == nil || info.IsDir() {
return nil
}
body, rerr := os.ReadFile(path)
if rerr != nil {
return nil
}
for label, secret := range map[string]string{"R": testR, "a wrong R": wrongR, "the repository password": repoPW} {
if strings.Contains(string(body), secret) {
t.Errorf("%s survived on disk at %s", label, path)
}
}
return nil
})
// And the staging directories are gone, not merely free of secrets.
entries, _ := os.ReadDir(tmp)
for _, e := range entries {
if e.IsDir() && strings.HasPrefix(e.Name(), "felhom-idesc-") {
t.Fatalf("an unseal staging directory survived: %s", e.Name())
}
}
}
// A fetch failure surfaces as a fetch failure, not as a wrong-code error — the operator must not be
// sent to re-read their recovery code because the hub was unreachable.
//
// ⚠ THIS TEST WAS GREEN THROUGHOUT THE DEFECT IT DESCRIBES (R-224, 2026-08-06). Its sentence is
// exactly right and it did not prevent anything, for two reasons worth keeping:
//
// 1. **It asserted the MECHANISM, one layer below the consequence.** It checked this package's error
// STRING. The merge happened one layer up, in the local-api handler's `default` branch, which
// answered a fetch failure with "the recovery code did not open the sealed bundle". The customer
// never sees this string; they see that one. The project's own rule — prefer the test that asserts
// the CONSEQUENCE (does the customer get blamed?) over the one that asserts the MECHANISM (is the
// error distinct here?) — names this case precisely.
// 2. **It asserted on TEXT.** `strings.Contains(err.Error(), …)` cannot be consumed by a caller, so
// it pinned something no production code could branch on. The distinction it checked was real and
// unusable.
//
// It now asserts the SENTINEL, which is what the handler branches on, and its consequence-level twin
// lives in `internal/localapi/escrow_recover_class_test.go` where the status is asserted.
func TestRecoverOffsiteRepoPassword_FetchErrorIsDistinct(t *testing.T) {
rec := OffsiteKeyRecoverer{Fetch: func(context.Context) ([]byte, bool, error) {
return nil, false, errors.New("hub: connection refused")
}}
_, err := rec.RecoverOffsiteRepoPassword(context.Background(), testR)
if err == nil || !errors.Is(err, ErrBundleFetch) {
t.Fatalf("a fetch failure must classify as ErrBundleFetch, got %v", err)
}
if errors.Is(err, ErrNoEscrowBlob) || errors.Is(err, ErrNoResticPassword) {
t.Fatal("a transport failure must not masquerade as a content verdict")
}
}
// ── R-224 — A FAILED FETCH IS NOT A WRONG CODE ──────────────────────────────────────────────────
//
// CAMPAIGN-11 F3 measured the consequence of these two being indistinguishable: with the hub
// firewalled off and a CORRECT current recovery code, the customer was told the code did not open
// their package, in 0.0556 s — no unseal was attempted at all.
//
// The pair below is the whole point. Asserting only the first would pass with a `return ErrBundleFetch`
// stuck on every error path, which is the same defect pointing the other way.
func TestRecoverOffsiteRepoPassword_FetchFailureIsClassifiedAsFetch(t *testing.T) {
boom := errors.New("hub: transport error: dial tcp 37.191.56.193:443: connect: no route to host")
r := OffsiteKeyRecoverer{Fetch: func(context.Context) ([]byte, bool, error) { return nil, false, boom }}
_, err := r.RecoverOffsiteRepoPassword(context.Background(), testR)
if err == nil {
t.Fatal("a failing fetch must return an error")
}
// RED-PROOF: drop the `%w: %w` join in RecoverOffsiteRepoPassword (return the bare wrapped cause,
// as it was before R-224) → this FAILS, and the local-api handler falls back to the wrong-code
// message exactly as it did on 2026-08-05.
if !errors.Is(err, ErrBundleFetch) {
t.Fatalf("a failed fetch must classify as ErrBundleFetch, got %v", err)
}
// The underlying cause survives for the operator log.
if !errors.Is(err, boom) {
t.Fatalf("the fetch cause must stay wrapped for the operator, got %v", err)
}
// And it must NOT be mistaken for either of the bundle-content situations.
if errors.Is(err, ErrNoEscrowBlob) || errors.Is(err, ErrNoResticPassword) {
t.Fatalf("a transport failure is neither of the bundle-content errors: %v", err)
}
}
// The other half: a genuinely wrong code must NOT classify as a fetch failure, or the fix trades one
// misattribution for its mirror image and the customer is told the hub is down when they mistyped.
func TestRecoverOffsiteRepoPassword_WrongCodeIsNotAFetchFailure(t *testing.T) {
ensureAge(t)
blob := sealBundle(t, IdentityBundle{ResticRepoPassword: "0123456789abcdef"}, testR)
r := OffsiteKeyRecoverer{Fetch: fetcherFor(blob)}
_, err := r.RecoverOffsiteRepoPassword(context.Background(),
"wrong horse battery staple sedative anaconda wobbly kingdom placard yodel")
if err == nil {
t.Fatal("a wrong recovery code must fail closed")
}
if errors.Is(err, ErrBundleFetch) {
t.Fatalf("a wrong code must NOT classify as a fetch failure, got %v", err)
}
}
// A clean "the hub holds nothing" keeps its own identity too — it is not a fetch failure, and the
// customer must not be told the hub was unreachable when it answered perfectly well.
func TestRecoverOffsiteRepoPassword_AbsentBlobIsNotAFetchFailure(t *testing.T) {
r := OffsiteKeyRecoverer{Fetch: func(context.Context) ([]byte, bool, error) { return nil, false, nil }}
_, err := r.RecoverOffsiteRepoPassword(context.Background(), testR)
if !errors.Is(err, ErrNoEscrowBlob) {
t.Fatalf("an absent blob must stay ErrNoEscrowBlob, got %v", err)
}
if errors.Is(err, ErrBundleFetch) {
t.Fatalf("an absent blob is not a fetch FAILURE, got %v", err)
}
}
+59
View File
@@ -0,0 +1,59 @@
// Package httpx holds the one HTTP-transport default this repo may not lose.
//
// Every client here pins TLS — PBS and PVE by leaf-cert SHA-256, the hub by an optional CA file —
// so none of them can use http.DefaultTransport and each hand-rolls its own. Hand-rolling silently
// discards DefaultTransport's settings, and one of them is load-bearing:
//
// Transport: &http.Transport{TLSClientConfig: tlsCfg} // IdleConnTimeout == 0 == NO timeout
//
// A zero IdleConnTimeout means idle keep-alive connections are retained FOREVER, not "use a sane
// default". Combined with a client that is rebuilt on a schedule and dropped (pbsTargetsFromPVE
// builds a fresh pbs.Client per cycle), every cycle strands one connection that nothing will ever
// close: the abandoned Transport becomes unreachable but its persistConn read-loop goroutine keeps
// the socket alive, and an unreachable Transport does not close its connections.
//
// Measured cost, live: 388 established connections accumulated on ep0's PBS proxy between
// 2026-08-18 09:51:22Z and 2026-08-20 08:02:13Z — 194 from each of the two boxes, held open on BOTH
// sides, one per agent poll cycle, on a proxy whose descriptor ceiling is 65536. See R-344 and
// felhom.eu/documentation/audits/SPIKE-ep0-established-connections-2026-08-20.md.
package httpx
import (
"crypto/tls"
"net/http"
"time"
)
// DefaultIdleConnTimeout is how long an idle keep-alive connection is retained before it is closed.
//
// It is 90s because that is http.DefaultTransport's own value: the fix for R-344 restores a
// standard-library default rather than inventing a number, so there is nothing here to tune and
// nothing to justify. It is comfortably shorter than every cadence that drives these clients (the
// 15-minute live-snapshot collect and the 6-hour verify loop), so a connection abandoned by one
// cycle is closed long before the next.
const DefaultIdleConnTimeout = 90 * time.Second
// NewTransport builds a FRESH *http.Transport pinned to tlsCfg, with the idle-connection timeout
// applied.
//
// Fresh, never shared: each caller pins a different endpoint, and a shared transport would pool
// connections across differently pinned servers. Reusing http.DefaultTransport for the same reason
// is not an option — it would drop the pin entirely.
//
// idleConnTimeout <= 0 means USE THE DEFAULT. It deliberately does not mean "no timeout": no-timeout
// is the bug this package exists to prevent, and an unset field must never be able to reintroduce
// it. Callers pass their configured value straight through; only tests pass a short one.
//
// Only IdleConnTimeout is set. The other DefaultTransport settings this transport also lacks
// (MaxIdleConns, TLSHandshakeTimeout, ExpectContinueTimeout) are deliberately left alone: none of
// them accumulates anything, every client bounds its whole request with http.Client.Timeout, and
// widening the change would have made the R-344 measurement unattributable.
func NewTransport(tlsCfg *tls.Config, idleConnTimeout time.Duration) *http.Transport {
if idleConnTimeout <= 0 {
idleConnTimeout = DefaultIdleConnTimeout
}
return &http.Transport{
TLSClientConfig: tlsCfg,
IdleConnTimeout: idleConnTimeout,
}
}
+74
View File
@@ -0,0 +1,74 @@
package httpx
import (
"crypto/tls"
"net/http"
"testing"
"time"
)
// TestNewTransport_ZeroMeansDefaultNeverForever is the whole point of this package.
//
// http.Transport's zero IdleConnTimeout means "retain idle connections FOREVER". Any code path that
// can reach that zero reintroduces R-344, so an unset, zero or negative value must all land on the
// default. If someone later "simplifies" NewTransport by passing the argument straight through,
// this fails.
func TestNewTransport_ZeroMeansDefaultNeverForever(t *testing.T) {
for _, tc := range []struct {
name string
in time.Duration
want time.Duration
}{
{"zero", 0, DefaultIdleConnTimeout},
{"negative", -time.Hour, DefaultIdleConnTimeout},
{"explicit short value (tests)", 50 * time.Millisecond, 50 * time.Millisecond},
{"explicit long value", time.Hour, time.Hour},
} {
t.Run(tc.name, func(t *testing.T) {
got := NewTransport(&tls.Config{MinVersion: tls.VersionTLS12}, tc.in).IdleConnTimeout
if got != tc.want {
t.Fatalf("IdleConnTimeout = %v, want %v", got, tc.want)
}
if got == 0 {
t.Fatal("IdleConnTimeout is 0 — that is 'never expire', which is the R-344 defect itself")
}
})
}
}
// TestDefaultIdleConnTimeout_MatchesTheStandardLibrary pins the number to its justification.
//
// 90s is not a tuned value; it is what http.DefaultTransport uses. Reading it off the standard
// library rather than hardcoding 90 means the constant cannot drift away from the reason given for
// it in the package doc.
func TestDefaultIdleConnTimeout_MatchesTheStandardLibrary(t *testing.T) {
std, ok := http.DefaultTransport.(*http.Transport)
if !ok {
t.Skip("http.DefaultTransport is not an *http.Transport in this Go build")
}
if DefaultIdleConnTimeout != std.IdleConnTimeout {
t.Fatalf("DefaultIdleConnTimeout = %v but http.DefaultTransport uses %v — the doc comment's justification no longer holds",
DefaultIdleConnTimeout, std.IdleConnTimeout)
}
}
// TestNewTransport_IsFreshEveryCall guards the pooling property the pinning relies on.
//
// Each caller pins a DIFFERENT endpoint. A shared transport would pool connections across
// differently pinned servers, so returning a package-level singleton would be a security change
// dressed as a tidy-up.
func TestNewTransport_IsFreshEveryCall(t *testing.T) {
a := NewTransport(&tls.Config{MinVersion: tls.VersionTLS12}, 0)
b := NewTransport(&tls.Config{MinVersion: tls.VersionTLS12}, 0)
if a == b {
t.Fatal("NewTransport returned the SAME transport twice — connections would be pooled across differently pinned endpoints")
}
}
// TestNewTransport_KeepsTheTLSConfig — the transport gains a field; it must lose nothing.
func TestNewTransport_KeepsTheTLSConfig(t *testing.T) {
cfg := &tls.Config{MinVersion: tls.VersionTLS12, InsecureSkipVerify: true} //nolint:gosec // test only
if got := NewTransport(cfg, 0).TLSClientConfig; got != cfg {
t.Fatalf("TLSClientConfig = %p, want the config passed in (%p) — the pin would be dropped", got, cfg)
}
}
+120 -2
View File
@@ -15,6 +15,8 @@ import (
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/config"
"gitea.dooplex.hu/admin/felhom-agent/internal/httpx"
)
const reportPath = "/api/v1/host-report"
@@ -49,8 +51,11 @@ func NewClient(cfg config.HubConfig, logger *slog.Logger) (*Client, error) {
tlsCfg.RootCAs = pool
}
hc := &http.Client{
Timeout: time.Duration(cfg.TimeoutSeconds) * time.Second,
Transport: &http.Transport{TLSClientConfig: tlsCfg},
Timeout: time.Duration(cfg.TimeoutSeconds) * time.Second,
// R-344, consistency only: this client is built ONCE per process, so it never accumulated
// and contributed nothing to the ep0 leak. It carried the same missing default, which over
// a tunnel is how one idle connection survives long enough to fail on next use.
Transport: httpx.NewTransport(tlsCfg, 0),
}
return newClient(cfg.URL, cfg.APIKey, cfg.HostID, hc, logger), nil
}
@@ -307,3 +312,116 @@ func tail(b []byte, max int) string {
}
return s
}
// IdentityEscrowResponse mirrors GET /api/v1/hosts/{host_id}/escrow (hub >= v0.94.0, R-199).
// Present=false is a CLEAN answer, not a fault: the host simply has no sealed bundle yet.
type IdentityEscrowResponse struct {
HostID string `json:"host_id"`
Present bool `json:"present"`
IdentityEscrowB64 string `json:"identity_escrow_b64"`
}
// FetchIdentityEscrow reads back THIS host's own opaque identity-escrow blob (R-199 link 6 — the
// mirror of UploadEscrow, self-scoped server-side by the per-host key). The bytes are ciphertext: they
// are useless without the customer's recovery code R, which neither the hub nor this agent ever holds.
//
// It is the ONLY retrieval this client performs, and it is deliberately narrow — no directive, no
// K-escrow, no key rotation. The operator-driven DR path (recovery-mode re-enroll) is a different
// endpoint with a different gate and is not reached from here.
//
// Errors are typed (transport vs HTTP) and never include the bearer token. The BLOB is never logged —
// only its length.
func (c *Client) FetchIdentityEscrow(ctx context.Context) (*IdentityEscrowResponse, error) {
if c.hostID == "" {
return nil, fmt.Errorf("hub: FetchIdentityEscrow requires a configured host_id")
}
url := c.baseURL + "/api/v1/hosts/" + c.hostID + "/escrow"
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
return nil, fmt.Errorf("hub: building escrow-fetch request: %w", err)
}
req.Header.Set("Authorization", "Bearer "+c.apiKey)
req.Header.Set("Accept", "application/json")
resp, err := c.hc.Do(req)
if err != nil {
return nil, &TransportError{Err: err}
}
defer resp.Body.Close()
raw, _ := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, &HTTPError{StatusCode: resp.StatusCode, BodyTail: tail(raw, 256)}
}
var out IdentityEscrowResponse
if err := json.Unmarshal(raw, &out); err != nil {
return nil, fmt.Errorf("hub: decoding escrow fetch: %w", err)
}
return &out, nil
}
// RetainedEscrowPackage is one RETAINED (superseded) sealed identity package. The blob is ciphertext
// and is useless without R. `SupersededAt` is the only thing here a human ever sees — it is what lets
// the recovery screen name WHICH earlier package a code belongs to.
type RetainedEscrowPackage struct {
Index int `json:"index"`
SupersededAt string `json:"superseded_at"`
KeyFingerprint string `json:"key_fingerprint"`
IdentityEscrowB64 string `json:"identity_escrow_b64"`
}
// RetainedEscrowResponse mirrors GET /api/v1/hosts/{host_id}/escrow/retained (hub >= v0.103.0, R-311).
//
// UnopenableCount is NOT noise. It counts retained packages the hub holds whose key material is absent
// (every pre-v0.93.0 row): on a box with those and nothing else, a perfectly correct old recovery code
// opens nothing, and the reason is a defect of ours. A caller that ignores this number will tell such a
// customer their code is wrong — the exact failure this whole chain exists to stop.
type RetainedEscrowResponse struct {
HostID string `json:"host_id"`
Count int `json:"count"`
UnopenableCount int `json:"unopenable_count"`
TruncatedCount int `json:"truncated_count"`
Packages []RetainedEscrowPackage `json:"packages"`
}
// FetchRetainedIdentityEscrow reads back THIS host's RETAINED sealed identity packages (R-311 —
// the retained siblings of FetchIdentityEscrow, self-scoped server-side by the same per-host key).
//
// SEPARATE FROM FetchIdentityEscrow ON PURPOSE. The ordinary recovery must not pay for this call, and
// must not fail because of it: the current package is tried first and alone, and this is reached only
// after that has refused. A hub too old to know this route answers 404, which is a CLEAN "none" here
// and must never be reported as a failed recovery.
func (c *Client) FetchRetainedIdentityEscrow(ctx context.Context) (*RetainedEscrowResponse, error) {
if c.hostID == "" {
return nil, fmt.Errorf("hub: FetchRetainedIdentityEscrow requires a configured host_id")
}
url := c.baseURL + "/api/v1/hosts/" + c.hostID + "/escrow/retained"
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
return nil, fmt.Errorf("hub: building retained-escrow request: %w", err)
}
req.Header.Set("Authorization", "Bearer "+c.apiKey)
req.Header.Set("Accept", "application/json")
resp, err := c.hc.Do(req)
if err != nil {
return nil, &TransportError{Err: err}
}
defer resp.Body.Close()
raw, _ := io.ReadAll(io.LimitReader(resp.Body, 4<<20))
if resp.StatusCode == http.StatusNotFound {
// A hub older than v0.103.0 has no such route. That is "no retained packages", not a fault —
// returning an error here would turn an old hub into a failed recovery on a box whose current
// package simply did not open.
return &RetainedEscrowResponse{HostID: c.hostID}, nil
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, &HTTPError{StatusCode: resp.StatusCode, BodyTail: tail(raw, 256)}
}
var out RetainedEscrowResponse
if err := json.Unmarshal(raw, &out); err != nil {
return nil, fmt.Errorf("hub: decoding retained escrow fetch: %w", err)
}
return &out, nil
}
+16
View File
@@ -103,6 +103,7 @@ type Collector struct {
addrEnum AddressEnumerator // v0.119.0: host interface enumeration; nil => the REAL one (see collectAddresses)
wg WireguardReporter // S3: offsite-tunnel status (nil → stanza omitted)
pbsdr PBSDRReporter // slice 2: PBS DR tier bridge state (nil → stanza omitted)
ctrlSup ControllerSupervisorReporter // R-523: in-guest controller supervisor (nil → stanza omitted)
guestNet GuestNetReporter // R-54: per-guest network watchdog (nil → stanza omitted)
selfUpdate SelfUpdateReporter // D1: agent self-update pending status (nil → false)
mgmtPlane MgmtPlaneReporter // G1: management-plane health (nil → stanza omitted)
@@ -195,6 +196,17 @@ func (c *Collector) SetPBSDRReporter(p PBSDRReporter) *Collector {
return c
}
// ControllerSupervisorReporter is the R-523 seam (satisfied by *localapi.Server).
type ControllerSupervisorReporter interface {
ControllerSupervisorStatus(ctx context.Context) *ControllerSupervisorStatus
}
// SetControllerSupervisorReporter wires the R-523 controller supervisor as a report source (nil-safe).
func (c *Collector) SetControllerSupervisorReporter(r ControllerSupervisorReporter) *Collector {
c.ctrlSup = r
return c
}
// SetGuestNetReporter wires the R-54 guest-network watchdog as a report source (nil-safe → stanza
// omitted). Returns the collector for chaining.
func (c *Collector) SetGuestNetReporter(g GuestNetReporter) *Collector {
@@ -285,6 +297,10 @@ func (c *Collector) Collect(ctx context.Context) (*HostReport, error) {
if c.pbsdr != nil {
report.PBSDR = c.pbsdr.PBSDRStatus(ctx)
}
// R-523: controller supervisor record (nil reporter = not wired → stanza omitted).
if c.ctrlSup != nil {
report.ControllerSupervisor = c.ctrlSup.ControllerSupervisorStatus(ctx)
}
// R-54: guest-network watchdog state (nil reporter = feature not wired → stanza omitted).
if c.guestNet != nil {
report.GuestNet = c.guestNet.GuestNetStatus(ctx)
+45 -2
View File
@@ -68,8 +68,14 @@ type HostReport struct {
// report is stored opaquely hub-side, so these additive fields need no hub-schema change.
// Both are `omitempty` (the Wireguard precedent): in the steady state (no update in flight)
// they are absent — which keeps the cross-repo host-report golden contract byte-stable without
// a hub change. They appear only while an update is pending. The hub reads an absent field as
// pending=false, the correct default.
// a hub change. They appear only while an update is pending.
//
// ⚠ CORRECTED 2026-08-08 (R-260). This comment used to end "The hub reads an absent field as
// pending=false, the correct default." THE HUB HAS NO FIELD FOR EITHER OF THESE, so it reads
// nothing — present or absent — and encoding/json discards them on arrival. The sentence
// described an intent, not the code, and it read as settled for long enough that a sweep had to
// find it. The emission is correct and stays; the missing consumer is tracked as R-264, and
// `felhom.eu/scripts/wire_contract_gate.py` now refuses any NEW field of this shape.
SelfUpdatePending bool `json:"selfupdate_pending,omitempty"`
SelfUpdatePendingVersion string `json:"selfupdate_pending_version,omitempty"`
@@ -118,6 +124,39 @@ type HostReport struct {
// HTTPS even when felhom-sshd or the tunnel is DOWN (channel independence). `omitempty`: absent
// when the feature is not wired (pre-H1) — additive, no hub-schema change.
OOB *OOBStatus `json:"oob,omitempty"`
// ControllerSupervisor (R-523, v0.131.0) is the in-guest controller supervisor's per-guest record:
// how many times the agent restarted a dead controller, when last and why, whether it gave up
// (crash-loop pause) and whether the operator parked it. The hub's ControllerSupervisorChecker
// mints `controller_restarted_by_agent` when last_restart_at MOVES and `controller_crashloop` when
// crashloop_since MOVES — timestamps, not counters, because the record is in-memory and an agent
// restart zeroes the counter. `omitempty`: absent when not wired, so the cross-repo golden stays
// byte-stable. The hub parser is pinned by hub/internal/monitor/controller_supervisor_test.go
// against the JSON TestControllerSupervisorStanza_WireShape pins here.
ControllerSupervisor *ControllerSupervisorStatus `json:"controller_supervisor,omitempty"`
}
// ControllerSupervisorStatus is the R-523 stanza. Carries no secret.
type ControllerSupervisorStatus struct {
Guests []ControllerSupervisorGuest `json:"guests"`
}
// ControllerSupervisorGuest is one supervised guest.
type ControllerSupervisorGuest struct {
VMID int `json:"vmid"`
RestartsTotal int `json:"restarts_total"`
LastRestartAt string `json:"last_restart_at,omitempty"` // RFC3339
LastReason string `json:"last_reason,omitempty"`
Crashloop bool `json:"crashloop"`
CrashloopSince string `json:"crashloop_since,omitempty"` // RFC3339; the last crash-loop, kept after it ends
Parked bool `json:"parked"`
// R-539 (v0.132.0) — the SLOW crash loop. Restarts24h counts restarts the supervisor performed in
// the last 24 hours (persisted, so an agent restart does not reset it); SlowCrashloop is true while
// the last raise is under 24 hours old; SlowCrashloopSince is the raise itself, which the hub keys on
// MOVING (hub v0.117.0 controller_slow_crashloop). It moves at most once per 24 hours.
Restarts24h int `json:"restarts_24h"`
SlowCrashloop bool `json:"slow_crashloop"`
SlowCrashloopSince string `json:"slow_crashloop_since,omitempty"` // RFC3339
}
// PBSDRStatus is the per-heartbeat PBS-DR-tier bridge state (slice 2). States:
@@ -431,6 +470,10 @@ type RestoreTest struct {
// mount layout, not just booted. Additive — a hub that predates them ignores the unknown keys.
MountParity string `json:"mount_parity,omitempty"`
MountInventory []string `json:"mount_inventory,omitempty"`
// Skipped (R-672, v0.133.0): the test did NOT run — the space preflight refused, and Error says
// why ("skipped: not enough space on …"). Pass is false. A hub that predates the key reads a
// failed test with that error, which is the honest reading.
Skipped bool `json:"skipped,omitempty"`
}
// PBSSnapshot is one PBS (offsite) snapshot's inventory + integrity state (doc 03 §8, slice
@@ -0,0 +1,133 @@
package localapi
import (
"context"
"encoding/json"
"io"
"log/slog"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// R-517 — the per-tier truth on GET /backup/status. BIGNIGHT: a successful 8.9 GB local backup,
// then a failed PBS attempt on a storage that did not exist; the page (fed by the single latest
// record) showed the 0-byte failure as "up to date" and the remote copy as present.
func tierStatesOf(t *testing.T, srv *Server) []TierBackupState {
t.Helper()
w := do(t, srv.Handler(), "GET", "/backup/status", "A", "")
var resp struct {
Data BackupStatusResponse `json:"data"`
}
if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
t.Fatalf("decode: %v (%s)", err, w.Body.String())
}
return resp.Data.Tiers
}
func tierStatesServer(t *testing.T, st *fakeStore, targets []hub.StorageTarget) *Server {
t.Helper()
srv, err := NewServer(Options{
ListenAddr: "127.0.0.1:0", Guests: &fakeGuests{}, Backups: &fakeBackups{}, Store: st,
Storage: fakeStorage{targets: targets},
Tokens: staticTokens{"A": 8200},
BackupTiers: []BackupTier{
{TargetID: "local", Cadence: 24 * time.Hour, Primary: true, Service: &fakeBackups{}},
{TargetID: "felhom-pbs", Cadence: 7 * 24 * time.Hour, Service: &fakeBackups{}},
},
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
})
if err != nil {
t.Fatal(err)
}
srv.baseCtx = context.Background()
srv.now = func() time.Time { return testNow }
return srv
}
// RED-PROOF (run 2026-09-15, recorded in REPORT.md): with tierBackupStates filling LastSuccess from
// pickLatestBackup(ctx, vmid, false, …) — an ATTEMPT — the pbs tier's last_success became the failed
// 0-byte record and this failed at "pbs tier reports a failed attempt as its last success".
func TestBackupStatus_TierStates_FailedTierNeverStandsInForSuccess(t *testing.T) {
st := &fakeStore{backups: []hub.Backup{
{TargetID: "local", VMID: 8200, Success: true, SizeBytes: 8877619753, StartedAt: "2026-06-10T11:03:23Z"},
{TargetID: "felhom-pbs", VMID: 8200, Success: false, Error: "storage 'felhom-pbs' does not exist", StartedAt: "2026-06-10T11:09:59Z"},
}}
srv := tierStatesServer(t, st, []hub.StorageTarget{{Name: "local", Type: "local"}}) // PBS storage ABSENT
tiers := tierStatesOf(t, srv)
if len(tiers) != 2 {
t.Fatalf("want 2 tiers, got %+v", tiers)
}
local, pbs := tiers[0], tiers[1]
if local.Target != "local" || local.LastSuccess == nil || local.LastSuccess.SizeBytes != 8877619753 || local.Storage != StoragePresencePresent {
t.Fatalf("local tier lost its successful backup: %+v", local)
}
if pbs.LastSuccess != nil {
t.Fatalf("pbs tier reports a failed attempt as its last success: %+v", pbs.LastSuccess)
}
if pbs.LastAttempt == nil || pbs.LastAttempt.Success || pbs.LastAttempt.Error == "" {
t.Fatalf("pbs tier's failed attempt is not reported as failed: %+v", pbs.LastAttempt)
}
if pbs.Storage != StoragePresenceAbsent {
t.Fatalf("pbs storage should read absent, got %q", pbs.Storage)
}
// The pre-R-517 field is unchanged (compat): still the newest record across targets.
w := do(t, srv.Handler(), "GET", "/backup/status", "A", "")
var resp struct {
Data BackupStatusResponse `json:"data"`
}
_ = json.Unmarshal(w.Body.Bytes(), &resp)
if resp.Data.Backup == nil || resp.Data.Backup.TargetID != "felhom-pbs" {
t.Fatalf("untargeted .backup changed meaning: %+v", resp.Data.Backup)
}
}
// /backup/tiers advertises storage presence, tri-state.
func TestBackupTiers_StoragePresence(t *testing.T) {
srv := tierStatesServer(t, &fakeStore{}, []hub.StorageTarget{{Name: "local"}})
w := do(t, srv.Handler(), "GET", "/backup/tiers", "A", "")
var resp struct {
Data BackupTiersResponse `json:"data"`
}
if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
t.Fatal(err)
}
got := map[string]string{}
for _, ti := range resp.Data.Tiers {
got[ti.Target] = ti.Storage
}
if got["local"] != "present" || got["felhom-pbs"] != "absent" {
t.Fatalf("storage presence wrong: %v", got)
}
// An unreadable storage view is "unknown", never "absent".
srv.storage = tierErrStorage{}
if p := srv.storagePresence(context.Background(), "felhom-pbs"); p != StoragePresenceUnknown {
t.Fatalf("unreadable storage view must be unknown, got %q", p)
}
}
type tierErrStorage struct{}
func (tierErrStorage) Observe(context.Context) ([]hub.StorageTarget, error) {
return nil, context.DeadlineExceeded
}
// A targeted request keeps the pre-R-517 bytes (no tiers array).
func TestBackupStatus_TargetedHasNoTiers(t *testing.T) {
srv := tierStatesServer(t, &fakeStore{}, []hub.StorageTarget{{Name: "local"}})
w := do(t, srv.Handler(), "GET", "/backup/status?target=local", "A", "")
if json.Valid(w.Body.Bytes()) && containsKey(w.Body.Bytes(), "tiers") {
t.Fatalf("targeted status grew a tiers array: %s", w.Body.String())
}
}
func containsKey(b []byte, key string) bool {
var m struct {
Data map[string]json.RawMessage `json:"data"`
}
_ = json.Unmarshal(b, &m)
_, ok := m.Data[key]
return ok
}
+453
View File
@@ -0,0 +1,453 @@
package localapi
import (
"context"
"encoding/json"
"os"
"path/filepath"
"sort"
"strconv"
"strings"
"sync"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// R-523 — the in-guest controller supervisor.
//
// THE OUTAGE THIS EXISTS TO KILL (BIGNIGHT F9, 2026-09-14). `docker kill felhom-controller` left the
// container `Exited (137)`. Nothing restarted it: Docker never restarts a container whose stop it
// records as deliberate — measured 2026-09-15 on Docker 29.8.0 for BOTH `unless-stopped` and `always`
// (evidence-p1fixes-2026-09-15/A1) — and the golden's `felhom-controller-bootstrap.service` is a
// oneshot (`RemainAfterExit=yes`) that ran once at boot and watches nothing. The household's
// dashboard answered 502 for 33 minutes until the box was power-cycled.
//
// This is doc 03 §4's sentence made real: "Healing a crashed controller is non-destructive by
// construction … redeploy = restart … inside the existing guest — never a guest destroy." The act
// is exactly the swap's own restart (`systemctl restart felhom-controller-bootstrap.service`, which
// does `docker rm -f` + `docker run` from the baked image and the guest's persistent volume), over
// the same GuestExecutor and the same two sudoers grants (`docker inspect -f *`, the unit restart).
// No new privilege.
//
// THE GUARDS, each because doing the act at the wrong moment is worse than not doing it:
// - not during a swap (the swap stops the controller ON PURPOSE and owns its own rollback);
// - not when the operator parked it (`<guests>/<vmid>/controller-parked` on the HOST);
// - not on a guest that is not running, is locked (backup/restore/snapshot/migrate), or has a
// vzdump in flight — a stopping or restoring guest is someone else's transaction;
// - not on ONE observation: the container must be seen not-running on two consecutive sweeps, so
// the bootstrap's own rm-f/run window (boot, path-unit hot-plug) is never raced;
// - no thrash: 3 restarts inside 15 minutes → stop restarting, raise `controller_crashloop`, try
// again after 30 minutes.
//
// THE EVENTS. The agent has no event channel of its own; its heartbeat IS the channel (the
// capability/leaf precedent). The per-guest record rides the host report as `controller_supervisor`,
// and the hub's ControllerSupervisorChecker mints `controller_restarted_by_agent` (info) when a
// guest's `last_restart_at` moves and `controller_crashloop` (error, operator-only) when
// `crashloop_since` moves. Timestamps, not counters, so an agent restart (which zeroes the in-memory
// record) can never read as a new restart.
const (
// controllerSupervisorInterval is the sweep cadence. Two not-running observations are required,
// so a killed controller is restarted 30–60 s after it died.
controllerSupervisorInterval = 30 * time.Second
// controllerSupervisorConfirm is how many consecutive not-running observations license a restart.
controllerSupervisorConfirm = 2
// Backoff: controllerCrashloopMax restarts inside controllerCrashloopWindow → give up for
// controllerCrashloopPause.
controllerCrashloopMax = 3
controllerCrashloopWindow = 15 * time.Minute
controllerCrashloopPause = 30 * time.Minute
// controllerSupervisorHeartbeatEvery: a liveness line every 20 sweeps (10 minutes) — a silent
// watchdog is indistinguishable from a dead one (standing rule 3).
controllerSupervisorHeartbeatEvery = 20
// ControllerParkedMarker is the host-side file that parks a guest's controller. The operator
// creates it with `touch /var/lib/felhom-agent/guests/<vmid>/controller-parked` and removes it to
// unpark. Host-side on purpose: it needs no in-guest exec grant, it survives a guest rebuild of
// the controller container, and a customer inside the guest cannot park the supervisor.
ControllerParkedMarker = "controller-parked"
defaultGuestsStateDir = "/var/lib/felhom-agent/guests"
// R-539 (operator ruling 3 of 2026-09-16) — the SLOW crash loop. The 3-in-15-minutes brake above
// cannot see a controller that dies every 20 minutes: no two restarts share its window, so it is
// restarted for ever and the only trace is an info event that mails nobody (measured 2026-09-16,
// R-531). A second counter over 24 hours raises a WARNING at the fifth restart. It does NOT stop
// restarting — the fast brake stays the only brake, unchanged. Every restart the supervisor
// performs counts, including one that follows a deliberate operator `docker kill` (measured
// 2026-09-15: the supervisor cannot tell a kill from a crash, and a controller that is killed five
// times a day is worth a line to the operator either way).
controllerSlowCrashloopWindow = 24 * time.Hour
controllerSlowCrashloopMax = 5
// controllerSlowCounterFile holds the 24-hour restart times and the last raise, per guest, beside
// the parked marker.
controllerSlowCounterFile = "controller-restarts-24h.json"
)
// controllerSupState is one guest's supervisor record. In-memory on purpose (the guest-power
// precedent): an agent restart forgets a crash-loop pause, which costs at most one more restart
// attempt, whereas persisting it could carry a stale "give up" across the restart that fixed it.
type controllerSupState struct {
notRunningSeen int
restarts []time.Time // restart times inside the crash-loop window (pruned)
restartsTotal int
lastRestartAt time.Time
lastReason string
crashloopSince time.Time // zero = not in a crash-loop pause
parked bool
// R-539 — the slow counter. PERSISTED, unlike everything above, and the precedent's reason does not
// apply to it: persisting the fast record could carry a stale "give up" across the restart that
// fixed it, but this record never gives anything up — it only warns. Losing it on an agent restart,
// on the other hand, would hide exactly the box it exists for (one whose agent restarts too).
restarts24h []time.Time
slowCrashloopSince time.Time // the last raise; kept after it ages out, the hub keys on it MOVING
}
type controllerSupervisor struct {
mu sync.Mutex
guests map[int]*controllerSupState
sweeps int
}
// WatchControllers runs the controller supervisor sweep until ctx is done. No-op when the guest list
// (staleLock) or the guest executor is not wired.
func (s *Server) WatchControllers(ctx context.Context) {
if s.staleLock == nil || s.guestExec == nil {
s.logger.Info("controller-supervisor: not wired (no guest list or no guest executor) — disabled")
return
}
s.logger.Info("controller-supervisor: started", "interval", controllerSupervisorInterval.String(),
"confirm_sweeps", controllerSupervisorConfirm, "crashloop_max", controllerCrashloopMax,
"crashloop_window", controllerCrashloopWindow.String(),
"slow_crashloop_max", controllerSlowCrashloopMax, "slow_crashloop_window", controllerSlowCrashloopWindow.String(),
"guests_dir", s.guestsStateDir())
t := time.NewTicker(controllerSupervisorInterval)
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case <-t.C:
s.ControllerSupervisorTick(ctx)
}
}
}
func (s *Server) guestsStateDir() string {
if s.guestsDir != "" {
return s.guestsDir
}
return defaultGuestsStateDir
}
// provisionedGuest reports whether the agent provisioned a controller into this guest: the
// `<guests>/<vmid>/bootstrap` directory exists. The directory itself, not bootstrap.json inside it —
// the directory is owned by the mapped guest root (0700), so the non-root agent can see the entry but
// not stat the file within.
func (s *Server) provisionedGuest(vmid int) bool {
fi, err := os.Stat(filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), "bootstrap"))
return err == nil && fi.IsDir()
}
func (s *Server) controllerParked(vmid int) bool {
_, err := os.Stat(filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), ControllerParkedMarker))
return err == nil
}
func (s *Server) supState(vmid int) *controllerSupState {
if s.ctrlSup.guests == nil {
s.ctrlSup.guests = map[int]*controllerSupState{}
}
st := s.ctrlSup.guests[vmid]
if st == nil {
st = &controllerSupState{}
s.loadSlowCounter(vmid, st)
s.ctrlSup.guests[vmid] = st
}
return st
}
// slowCounterRecord is the on-disk shape of the R-539 counter.
type slowCounterRecord struct {
Restarts []time.Time `json:"restarts"`
SlowCrashloopSince time.Time `json:"slow_crashloop_since,omitempty"`
}
func (s *Server) slowCounterPath(vmid int) string {
return filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), controllerSlowCounterFile)
}
// loadSlowCounter restores the persisted counter into a fresh state. Absent = a clean start; unreadable
// or corrupt = a clean start with a WARN (a warning counter must never block supervision).
func (s *Server) loadSlowCounter(vmid int, st *controllerSupState) {
b, err := os.ReadFile(s.slowCounterPath(vmid))
if err != nil {
if !os.IsNotExist(err) {
s.logger.Warn("controller-supervisor: slow counter unreadable — starting it from zero", "vmid", vmid, "err", err)
}
return
}
var rec slowCounterRecord
if err := json.Unmarshal(b, &rec); err != nil {
s.logger.Warn("controller-supervisor: slow counter corrupt — starting it from zero", "vmid", vmid, "err", err)
return
}
st.restarts24h = pruneBefore(rec.Restarts, s.clock().Add(-controllerSlowCrashloopWindow))
st.slowCrashloopSince = rec.SlowCrashloopSince
if len(st.restarts24h) > 0 || !st.slowCrashloopSince.IsZero() {
s.logger.Info("controller-supervisor: slow counter restored from disk", "vmid", vmid,
"restarts_24h", len(st.restarts24h), "slow_crashloop_since", st.slowCrashloopSince.Format(time.RFC3339))
}
}
// saveSlowCounter writes the counter atomically (tmp + rename, 0600). A failure is logged and the
// in-memory counter carries on — the next restart retries the write.
func (s *Server) saveSlowCounter(vmid int, rec slowCounterRecord) {
path := s.slowCounterPath(vmid)
b, err := json.Marshal(rec)
if err == nil {
tmp := path + ".tmp"
if err = os.WriteFile(tmp, b, 0o600); err == nil {
err = os.Rename(tmp, path)
}
}
if err != nil {
s.logger.Warn("controller-supervisor: could not persist the slow counter (kept in memory)", "vmid", vmid, "path", path, "err", err)
}
}
// ControllerSupervisorTick performs one sweep. Exported so a test (and a live check) can drive one
// cycle without waiting on the ticker.
func (s *Server) ControllerSupervisorTick(ctx context.Context) {
if s.staleLock == nil || s.guestExec == nil {
return
}
guests, err := s.staleLock.Guests(ctx)
if err != nil {
// Ownership unproven ⇒ touch nothing (the guest-power rule).
s.logger.Warn("controller-supervisor: guest list unavailable — skipping sweep (ownership unproven)", "err", err)
return
}
var evaluated, down int
for _, g := range guests {
if ctx.Err() != nil {
return
}
if !s.provisionedGuest(g.VMID) {
continue
}
evaluated++
if !s.superviseOneController(ctx, g.VMID, g.Status) {
down++
}
}
s.ctrlSup.mu.Lock()
s.ctrlSup.sweeps++
sweeps := s.ctrlSup.sweeps
s.ctrlSup.mu.Unlock()
if sweeps%controllerSupervisorHeartbeatEvery == 0 {
s.logger.Info("controller-supervisor: alive", "sweeps_since_boot", sweeps,
"guests_evaluated", evaluated, "controllers_not_running", down)
}
}
// controllerRunning asks the guest's Docker for the controller's state. Returns (running, known).
// known=false means the question could not be answered (pct exec failed for a reason other than a
// missing container) — the caller does nothing on unknown. An ABSENT container is a known "not
// running": `docker rm` of the controller is the same outage as a kill.
func (s *Server) controllerRunning(ctx context.Context, vmid int) (running, known bool, status string) {
out, err := s.guestExec.GuestExec(ctx, vmid, "docker", "inspect", "-f", "{{.State.Status}}", controllerContainer)
if err != nil {
msg := strings.ToLower(err.Error() + " " + out)
if strings.Contains(msg, "no such object") || strings.Contains(msg, "no such container") {
return false, true, "absent"
}
return false, false, ""
}
status = strings.TrimSpace(out)
// "restarting" is Docker's own restart loop at work — not ours to fight on this sweep.
return status == "running" || status == "restarting", true, status
}
// superviseOneController evaluates one provisioned guest and restarts its controller when every guard
// allows. Returns false when the controller was observed not running.
func (s *Server) superviseOneController(ctx context.Context, vmid int, guestStatus string) bool {
now := s.clock()
if guestStatus != "running" {
s.resetNotRunning(vmid)
return true // the guest-power watchdog owns a stopped guest; its controller is not "down"
}
running, known, status := s.controllerRunning(ctx, vmid)
if !known {
s.logger.Debug("controller-supervisor: controller state unknown (guest exec failed) — no action", "vmid", vmid)
s.resetNotRunning(vmid)
return true
}
parked := s.controllerParked(vmid)
s.ctrlSup.mu.Lock()
st := s.supState(vmid)
st.parked = parked
if running {
st.notRunningSeen = 0
s.ctrlSup.mu.Unlock()
return true
}
st.notRunningSeen++
seen := st.notRunningSeen
s.ctrlSup.mu.Unlock()
if parked {
s.logger.Info("controller-supervisor: controller is not running and the guest is PARKED — leaving it",
"vmid", vmid, "status", status, "marker", filepath.Join(s.guestsStateDir(), strconv.Itoa(vmid), ControllerParkedMarker))
return false
}
s.swapMu.Lock()
swapping := s.swapInFlight[vmid]
s.swapMu.Unlock()
if swapping {
s.logger.Info("controller-supervisor: controller is not running during a controller SWAP — the swap owns it",
"vmid", vmid, "status", status)
s.resetNotRunning(vmid)
return false
}
if seen < controllerSupervisorConfirm {
s.logger.Info("controller-supervisor: controller observed not running — confirming on the next sweep",
"vmid", vmid, "status", status, "seen", seen, "of", controllerSupervisorConfirm)
return false
}
lock, _, err := s.staleLock.Lock(ctx, vmid)
if err != nil {
s.logger.Warn("controller-supervisor: could not read the guest lock — no action (fail-safe)", "vmid", vmid, "err", err)
return false
}
if lock != "" {
s.logger.Info("controller-supervisor: guest is LOCKED — another operation owns it, no action", "vmid", vmid, "lock", lock)
return false
}
if busy, berr := s.staleLock.BackupRunning(ctx, vmid); berr != nil || busy {
s.logger.Info("controller-supervisor: a vzdump may be in flight for the guest — no action",
"vmid", vmid, "backup_running", busy, "err", berr)
return false
}
// Backoff.
s.ctrlSup.mu.Lock()
st = s.supState(vmid)
if !st.crashloopSince.IsZero() {
if now.Sub(st.crashloopSince) < controllerCrashloopPause {
s.ctrlSup.mu.Unlock()
s.logger.Warn("controller-supervisor: crash-loop pause in force — not restarting",
"vmid", vmid, "since", st.crashloopSince.Format(time.RFC3339), "resume_after", controllerCrashloopPause.String())
return false
}
// Pause over: resume with a clean window. crashloopSince stays as the record of the last
// crash-loop (the hub keys on it moving, not on it clearing).
st.restarts = nil
st.crashloopSince = time.Time{}
}
st.restarts = pruneBefore(st.restarts, now.Add(-controllerCrashloopWindow))
if len(st.restarts) >= controllerCrashloopMax {
st.crashloopSince = now
n := len(st.restarts)
s.ctrlSup.mu.Unlock()
s.logger.Error("controller-supervisor: CRASH-LOOP — the controller would not stay up; stopping restarts and raising controller_crashloop",
"vmid", vmid, "restarts_in_window", n, "window", controllerCrashloopWindow.String(), "pause", controllerCrashloopPause.String())
return false
}
s.ctrlSup.mu.Unlock()
reason := "controller container " + status + " on " + strconv.Itoa(controllerSupervisorConfirm) + " consecutive sweeps"
s.logger.Warn("controller-supervisor: controller is NOT running — restarting the bootstrap unit",
"vmid", vmid, "status", status, "unit", bootstrapUnit)
if _, err := s.guestExec.GuestExec(ctx, vmid, "systemctl", "restart", bootstrapUnit); err != nil {
s.logger.Error("controller-supervisor: bootstrap restart failed", "vmid", vmid, "err", err)
reason += "; restart FAILED: " + err.Error()
}
s.ctrlSup.mu.Lock()
st = s.supState(vmid)
st.restarts = append(st.restarts, now)
st.restartsTotal++
st.lastRestartAt = now
st.lastReason = reason
st.notRunningSeen = 0
// R-539: the slow counter. Raise at most once per 24 hours — the hub mails on the raise MOVING.
st.restarts24h = append(pruneBefore(st.restarts24h, now.Add(-controllerSlowCrashloopWindow)), now)
n24 := len(st.restarts24h)
raised := false
if n24 >= controllerSlowCrashloopMax && (st.slowCrashloopSince.IsZero() || now.Sub(st.slowCrashloopSince) >= controllerSlowCrashloopWindow) {
st.slowCrashloopSince = now
raised = true
}
rec := slowCounterRecord{Restarts: append([]time.Time(nil), st.restarts24h...), SlowCrashloopSince: st.slowCrashloopSince}
s.ctrlSup.mu.Unlock()
s.saveSlowCounter(vmid, rec)
s.logger.Warn("controller-supervisor: RESTARTED the controller", "vmid", vmid, "reason", reason, "restarts_24h", n24)
if raised {
s.logger.Warn("controller-supervisor: SLOW CRASH-LOOP — the controller keeps dying; still restarting it, raising controller_slow_crashloop",
"vmid", vmid, "restarts_24h", n24, "window", controllerSlowCrashloopWindow.String(), "threshold", controllerSlowCrashloopMax)
}
return false
}
func (s *Server) resetNotRunning(vmid int) {
s.ctrlSup.mu.Lock()
defer s.ctrlSup.mu.Unlock()
if st := s.ctrlSup.guests[vmid]; st != nil {
st.notRunningSeen = 0
}
}
func (s *Server) clock() time.Time {
if s.now != nil {
return s.now()
}
return time.Now().UTC()
}
func pruneBefore(ts []time.Time, cutoff time.Time) []time.Time {
out := ts[:0]
for _, t := range ts {
if !t.Before(cutoff) {
out = append(out, t)
}
}
return out
}
// ControllerSupervisorStatus is the host-report stanza source (hub.ControllerSupervisorReporter).
// Nil when the supervisor is not wired, so the stanza is omitted.
func (s *Server) ControllerSupervisorStatus(_ context.Context) *hub.ControllerSupervisorStatus {
if s.staleLock == nil || s.guestExec == nil {
return nil
}
s.ctrlSup.mu.Lock()
defer s.ctrlSup.mu.Unlock()
out := &hub.ControllerSupervisorStatus{Guests: []hub.ControllerSupervisorGuest{}}
for vmid, st := range s.ctrlSup.guests {
g := hub.ControllerSupervisorGuest{
VMID: vmid,
RestartsTotal: st.restartsTotal,
LastReason: st.lastReason,
Parked: st.parked,
Crashloop: !st.crashloopSince.IsZero(),
}
if !st.lastRestartAt.IsZero() {
g.LastRestartAt = st.lastRestartAt.UTC().Format(time.RFC3339)
}
if !st.crashloopSince.IsZero() {
g.CrashloopSince = st.crashloopSince.UTC().Format(time.RFC3339)
}
now := s.clock()
g.Restarts24h = len(pruneBefore(append([]time.Time(nil), st.restarts24h...), now.Add(-controllerSlowCrashloopWindow)))
if !st.slowCrashloopSince.IsZero() {
g.SlowCrashloopSince = st.slowCrashloopSince.UTC().Format(time.RFC3339)
g.SlowCrashloop = now.Sub(st.slowCrashloopSince) < controllerSlowCrashloopWindow
}
out.Guests = append(out.Guests, g)
}
sort.Slice(out.Guests, func(i, j int) bool { return out.Guests[i].VMID < out.Guests[j].VMID })
return out
}
@@ -0,0 +1,350 @@
package localapi
import (
"context"
"encoding/json"
"errors"
"io"
"log/slog"
"os"
"path/filepath"
"strconv"
"sync"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-523 — the controller supervisor. The consequence under test is "a dead controller is started
// again", and each guard is pinned by the case where acting would be wrong.
type supExec struct {
mu sync.Mutex
status map[int]string // docker .State.Status per vmid; "" = container absent
inspectErr error // non-nil = pct exec itself failed (unknown)
restarts map[int]int
// onRestart, when set, is the status the container reaches after a restart (a crash-looper
// stays "exited").
onRestart string
}
func (f *supExec) GuestExec(_ context.Context, vmid int, args ...string) (string, error) {
f.mu.Lock()
defer f.mu.Unlock()
switch {
case len(args) >= 5 && args[0] == "docker" && args[1] == "inspect" && args[3] == "{{.State.Status}}":
if f.inspectErr != nil {
return "", f.inspectErr
}
st, ok := f.status[vmid]
if !ok || st == "" {
return "", errors.New("pct exec: exit status 1: Error: No such object: felhom-controller")
}
return st + "\n", nil
case len(args) == 3 && args[0] == "systemctl" && args[1] == "restart" && args[2] == bootstrapUnit:
if f.restarts == nil {
f.restarts = map[int]int{}
}
f.restarts[vmid]++
if f.onRestart != "" {
f.status[vmid] = f.onRestart
}
return "", nil
}
return "", errors.New("supExec: unexpected args")
}
func (f *supExec) GuestExecStdin(context.Context, int, io.Reader, ...string) (string, error) {
return "", errors.New("supExec: no stdin exec expected")
}
func (f *supExec) count(vmid int) int {
f.mu.Lock()
defer f.mu.Unlock()
return f.restarts[vmid]
}
type supClock struct{ t time.Time }
func (c *supClock) now() time.Time { return c.t }
func supServer(t *testing.T, ex *supExec, ctl *fakeGuestPowerCtl, provisioned ...int) (*Server, *supClock, string) {
t.Helper()
dir := t.TempDir()
for _, v := range provisioned {
if err := os.MkdirAll(filepath.Join(dir, strconv.Itoa(v), "bootstrap"), 0o700); err != nil {
t.Fatal(err)
}
}
clk := &supClock{t: time.Date(2026, 9, 15, 10, 0, 0, 0, time.UTC)}
s := &Server{
staleLock: ctl,
guestExec: ex,
guestsDir: dir,
swapInFlight: map[int]bool{},
logger: slog.New(slog.NewTextHandler(discardW{}, nil)),
now: clk.now,
}
return s, clk, dir
}
func runningGuest(vmid int) *fakeGuestPowerCtl {
return &fakeGuestPowerCtl{
guests: []proxmox.Guest{{VMID: vmid, Status: "running"}},
locks: map[int]string{}, onboot: map[int]bool{vmid: true},
}
}
// The consequence: a killed controller IS restarted — on the second consecutive observation, not
// the first (the bootstrap's own rm-f/run window must never be raced).
//
// RED-PROOF: delete the `systemctl restart` GuestExec call in superviseOneController → restarts
// stays 0 → "the killed controller was NOT restarted — this is R-523".
func TestControllerSupervisor_KilledControllerIsRestarted(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "running"}
s, _, _ := supServer(t, ex, runningGuest(9201), 9201)
s.ControllerSupervisorTick(context.Background())
if n := ex.count(9201); n != 0 {
t.Fatalf("restarted on the FIRST observation (restarts=%d) — must confirm on a second sweep", n)
}
s.ControllerSupervisorTick(context.Background())
if n := ex.count(9201); n != 1 {
t.Fatalf("the killed controller was NOT restarted — this is R-523 (restarts=%d)", n)
}
st := s.ControllerSupervisorStatus(context.Background())
if len(st.Guests) != 1 || st.Guests[0].RestartsTotal != 1 || st.Guests[0].LastRestartAt == "" || st.Guests[0].LastReason == "" {
t.Fatalf("report stanza did not record the restart: %+v", st.Guests)
}
// Healthy again → no further restart.
s.ControllerSupervisorTick(context.Background())
s.ControllerSupervisorTick(context.Background())
if n := ex.count(9201); n != 1 {
t.Fatalf("a running controller was restarted again (restarts=%d)", n)
}
}
func TestControllerSupervisor_AbsentContainerIsRestarted(t *testing.T) {
ex := &supExec{status: map[int]string{}, onRestart: "running"}
s, _, _ := supServer(t, ex, runningGuest(9201), 9201)
s.ControllerSupervisorTick(context.Background())
s.ControllerSupervisorTick(context.Background())
if n := ex.count(9201); n != 1 {
t.Fatalf("a removed controller container was not restarted (restarts=%d)", n)
}
}
func TestControllerSupervisor_Guards(t *testing.T) {
cases := []struct {
name string
setup func(s *Server, ex *supExec, ctl *fakeGuestPowerCtl, dir string)
}{
{"parked", func(s *Server, _ *supExec, _ *fakeGuestPowerCtl, dir string) {
if err := os.WriteFile(filepath.Join(dir, "9201", ControllerParkedMarker), nil, 0o600); err != nil {
panic(err)
}
}},
{"swap in flight", func(s *Server, _ *supExec, _ *fakeGuestPowerCtl, _ string) { s.swapInFlight[9201] = true }},
{"guest locked", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) { ctl.locks[9201] = "backup" }},
{"vzdump running", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
ctl.backupRun = map[int]bool{9201: true}
}},
{"vzdump state unknown", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
ctl.backupErr = errors.New("tasks unreadable")
}},
{"guest not running", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
ctl.guests[0].Status = "stopped"
}},
{"docker state unknown", func(_ *Server, ex *supExec, _ *fakeGuestPowerCtl, _ string) {
ex.inspectErr = errors.New("pct exec 9201: exit status 255: container not running")
}},
{"guest list unavailable", func(_ *Server, _ *supExec, ctl *fakeGuestPowerCtl, _ string) {
ctl.guestsErr = errors.New("api down")
}},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "running"}
ctl := runningGuest(9201)
s, _, dir := supServer(t, ex, ctl, 9201)
tc.setup(s, ex, ctl, dir)
for i := 0; i < 4; i++ {
s.ControllerSupervisorTick(context.Background())
}
if n := ex.count(9201); n != 0 {
t.Fatalf("guard %q did not hold: controller restarted %d time(s)", tc.name, n)
}
})
}
}
// A guest the agent did not provision (no <guests>/<vmid>/bootstrap) is never touched.
func TestControllerSupervisor_UnprovisionedGuestIgnored(t *testing.T) {
ex := &supExec{status: map[int]string{9202: "exited"}}
s, _, _ := supServer(t, ex, runningGuest(9202) /* nothing provisioned */)
for i := 0; i < 3; i++ {
s.ControllerSupervisorTick(context.Background())
}
if n := ex.count(9202); n != 0 {
t.Fatalf("an unprovisioned guest's container was restarted (%d)", n)
}
}
// No thrash: a controller that will not stay up is restarted at most controllerCrashloopMax times
// inside the window, then the supervisor raises the crash-loop and pauses; after the pause it tries
// again.
//
// RED-PROOF (run 2026-09-15, recorded in REPORT.md): with the `len(st.restarts) >=
// controllerCrashloopMax` block removed, restarts reached 10 in the first 20 sweeps and the test
// failed at "crash-looping controller restarted 10 times".
func TestControllerSupervisor_CrashloopBackoff(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "exited"}
s, clk, _ := supServer(t, ex, runningGuest(9201), 9201)
ctx := context.Background()
for i := 0; i < 20; i++ { // 10 minutes of 30 s sweeps
s.ControllerSupervisorTick(ctx)
clk.t = clk.t.Add(controllerSupervisorInterval)
}
if n := ex.count(9201); n != controllerCrashloopMax {
t.Fatalf("crash-looping controller restarted %d times in 10 minutes — want exactly %d then a pause", n, controllerCrashloopMax)
}
st := s.ControllerSupervisorStatus(ctx).Guests[0]
if !st.Crashloop || st.CrashloopSince == "" {
t.Fatalf("crash-loop not raised in the report stanza: %+v", st)
}
// Still paused 25 minutes later.
clk.t = clk.t.Add(15 * time.Minute)
s.ControllerSupervisorTick(ctx)
s.ControllerSupervisorTick(ctx)
if n := ex.count(9201); n != controllerCrashloopMax {
t.Fatalf("restarted during the crash-loop pause (restarts=%d)", n)
}
// After the pause: tries again.
clk.t = clk.t.Add(controllerCrashloopPause)
s.ControllerSupervisorTick(ctx)
s.ControllerSupervisorTick(ctx)
if n := ex.count(9201); n != controllerCrashloopMax+1 {
t.Fatalf("did not resume after the pause (restarts=%d, want %d)", n, controllerCrashloopMax+1)
}
if since := s.ControllerSupervisorStatus(ctx).Guests[0].CrashloopSince; since != st.CrashloopSince && since != "" {
t.Fatalf("crashloop_since changed without a new crash-loop: %q → %q", st.CrashloopSince, since)
}
}
// The wire shape the hub parses. The hub's controller_supervisor_test.go carries this exact JSON.
func TestControllerSupervisorStanza_WireShape(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "exited"}, onRestart: "running"}
s, _, _ := supServer(t, ex, runningGuest(9201), 9201)
s.ControllerSupervisorTick(context.Background())
s.ControllerSupervisorTick(context.Background())
b, _ := json.Marshal(s.ControllerSupervisorStatus(context.Background()))
var m map[string][]map[string]any
if err := json.Unmarshal(b, &m); err != nil {
t.Fatal(err)
}
g := m["guests"][0]
for _, k := range []string{"vmid", "restarts_total", "last_restart_at", "last_reason", "crashloop", "parked", "restarts_24h", "slow_crashloop"} {
if _, ok := g[k]; !ok {
t.Fatalf("stanza lacks %q — the hub keys on it: %s", k, b)
}
}
}
// ---- R-539 (operator ruling 3 of 2026-09-16): the SLOW crash loop ----------------------------------
// supKillOnce kills the controller and lets the supervisor restart it (two confirming sweeps), then
// moves the clock on by gap. The container comes back "running", so each restart is a separate act.
func supKillOnce(t *testing.T, s *Server, ex *supExec, clk *supClock, gap time.Duration) {
t.Helper()
before := ex.count(9201)
ex.mu.Lock()
ex.status[9201] = "exited"
ex.mu.Unlock()
s.ControllerSupervisorTick(context.Background())
clk.t = clk.t.Add(controllerSupervisorInterval)
s.ControllerSupervisorTick(context.Background())
if ex.count(9201) != before+1 {
t.Fatalf("kill was not followed by exactly one restart (restarts %d → %d)", before, ex.count(9201))
}
clk.t = clk.t.Add(gap)
}
// The consequence: a controller that dies every 20 minutes — never three times inside the 15-minute
// brake — raises slow_crashloop on the FIFTH restart in 24 hours, and the raise does not move again on
// the sixth (the hub mails on movement; once per 24 hours is the ruling).
//
// RED-PROOF: without the slow counter the stanza never sets slow_crashloop → "five restarts 20 minutes
// apart did not raise slow_crashloop — this is R-539".
func TestControllerSupervisor_SlowCrashloop(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "running"}, onRestart: "running"}
s, clk, _ := supServer(t, ex, runningGuest(9201), 9201)
ctx := context.Background()
for i := 1; i <= 4; i++ {
supKillOnce(t, s, ex, clk, 20*time.Minute)
}
g := s.ControllerSupervisorStatus(ctx).Guests[0]
if g.Crashloop {
t.Fatalf("the 15-minute brake fired on restarts 20 minutes apart — the fixture is wrong: %+v", g)
}
if g.SlowCrashloop || g.SlowCrashloopSince != "" {
t.Fatalf("slow_crashloop raised after only 4 restarts: %+v", g)
}
supKillOnce(t, s, ex, clk, 20*time.Minute)
g = s.ControllerSupervisorStatus(ctx).Guests[0]
if !g.SlowCrashloop || g.SlowCrashloopSince == "" || g.Restarts24h != 5 {
t.Fatalf("five restarts 20 minutes apart did not raise slow_crashloop — this is R-539: %+v", g)
}
first := g.SlowCrashloopSince
supKillOnce(t, s, ex, clk, 20*time.Minute)
g = s.ControllerSupervisorStatus(ctx).Guests[0]
if g.SlowCrashloopSince != first {
t.Fatalf("the raise moved again on the 6th restart (%q → %q) — the operator would be mailed per restart", first, g.SlowCrashloopSince)
}
if !g.SlowCrashloop {
t.Fatalf("slow_crashloop cleared while the loop continues: %+v", g)
}
}
// The negative control: restarts that never reach five inside any 24 hours never raise it.
func TestControllerSupervisor_SpreadRestartsNeverSlowCrashloop(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "running"}, onRestart: "running"}
s, clk, _ := supServer(t, ex, runningGuest(9201), 9201)
for i := 0; i < 8; i++ { // eight restarts, 7 hours apart: at most 4 inside any 24 hours
supKillOnce(t, s, ex, clk, 7*time.Hour)
}
g := s.ControllerSupervisorStatus(context.Background()).Guests[0]
if g.SlowCrashloop || g.SlowCrashloopSince != "" {
t.Fatalf("restarts 7 hours apart raised slow_crashloop: %+v", g)
}
if g.Restarts24h > 4 {
t.Fatalf("restarts_24h=%d — the 24-hour window is not pruning", g.Restarts24h)
}
}
// An agent restart must not reset the slow counter (the ruling; a box whose AGENT also restarts would
// otherwise never reach five). The same state directory, a fresh Server.
//
// RED-PROOF: keep the counter in memory only → the second Server starts at 0 → "the agent restart
// reset the slow counter".
func TestControllerSupervisor_SlowCounterSurvivesAgentRestart(t *testing.T) {
ex := &supExec{status: map[int]string{9201: "running"}, onRestart: "running"}
s, clk, dir := supServer(t, ex, runningGuest(9201), 9201)
for i := 0; i < 4; i++ {
supKillOnce(t, s, ex, clk, 20*time.Minute)
}
s2 := &Server{
staleLock: s.staleLock,
guestExec: ex,
guestsDir: dir,
swapInFlight: map[int]bool{},
logger: s.logger,
now: clk.now,
}
supKillOnce(t, s2, ex, clk, 20*time.Minute)
g := s2.ControllerSupervisorStatus(context.Background()).Guests[0]
if g.Restarts24h != 5 || !g.SlowCrashloop {
t.Fatalf("the agent restart reset the slow counter: %+v", g)
}
}
+153
View File
@@ -0,0 +1,153 @@
package localapi
import (
"context"
"crypto/sha256"
"encoding/hex"
"errors"
"net/http"
"strings"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/escrow"
)
// R-199 (agent v0.125.0) — the in-guest controller asks the agent to recover the offsite repository
// password from the hub's sealed bundle, using the customer's recovery code R.
//
// WHY THE AGENT AND NOT THE CONTROLLER. Three reasons, all structural: the unsealing binary (`age`)
// is an agent runtime dependency and is deliberately absent from the controller image; the sealed
// blob is a HOST-scoped object whose only writer is this agent under the per-host key, so the read is
// that write's mirror; and the controller is a trust tier down — it should receive one field, not a
// bundle it has no use for.
//
// R'S HANDLING, WHICH IS THE TIGHTEST RULE IN THIS FLOW. R is the one secret in the system that
// cannot be rotated, re-issued or recovered — it exists only in the customer's hands. Here it:
// - arrives in the request body over the already-pinned local-API channel (the operator accepted
// that crossing on 2026-08-04; the acceptance covers the CHANNEL, not carelessness at either end);
// - is held in memory for the duration of one call and cleared on BOTH paths;
// - is never written to disk, never an argument in a process list, and never logged at any level,
// including inside an error;
// - is never echoed: no response this endpoint can emit contains it.
//
// The request-level DEBUG middleware logs method/path/status/duration and never bodies — see
// `logRequests`. Do not add a body dump.
//
// THE RESPONSE CARRIES THE PASSWORD AND ITS HASH. The hash is what this session's proof compares
// (compare by hash, never by value). The password itself is present because the next link — placing a
// recovered password so the existing repository opens — needs it, and building a hash-only seam now
// would have to be torn out to add it. The controller's diagnostic reads only the hash.
type recoverOffsitePasswordRequest struct {
VMID int `json:"vmid"`
// RecoveryCode is the customer's R. NEVER logged, never persisted, never echoed.
RecoveryCode string `json:"recovery_code"`
}
// handleRecoverOffsitePassword fetches this host's sealed bundle, unseals it with R and returns only
// the offsite repository password (plus its sha256, for hash-only comparison by the caller).
func (s *Server) handleRecoverOffsitePassword(w http.ResponseWriter, r *http.Request, vmid int) {
var req recoverOffsitePasswordRequest
if !decodeBody(w, r, &req) {
return
}
if !s.scopedFromBody(w, req.VMID, vmid, r.URL.Path) {
return
}
R := strings.TrimSpace(req.RecoveryCode)
req.RecoveryCode = "" // drop the decoded copy immediately
if R == "" {
writeErr(w, http.StatusBadRequest, "recovery_code is required")
return
}
if s.escrowRecovery == nil {
R = ""
writeErr(w, http.StatusServiceUnavailable, "offsite key recovery is not configured on this agent (no hub client)")
return
}
ctx, cancel := context.WithTimeout(r.Context(), 60*time.Second)
defer cancel()
s.logger.Info("local-api: recovering the offsite repository password from the sealed escrow (R via body, never logged/persisted)", "vmid", vmid)
pw, err := s.escrowRecovery.RecoverOffsiteRepoPassword(ctx, R)
R = "" // cleared on BOTH paths, before anything else can happen
if err != nil {
// Each situation gets its own status and its own words. None of them names a secret.
switch {
// ── R-224 (2026-08-06) — THE FETCH FAILURE IS NOT A WRONG CODE. ────────────────────────
//
// This case did not exist, and its absence is the defect. A failed fetch fell through to the
// `default` below and was answered with "the recovery code did not open the sealed bundle" —
// so a hub that could not be reached was reported to the customer as a bad recovery code, on
// the one screen whose whole purpose is to be believed about their backups.
//
// Measured live 2026-08-05 (CAMPAIGN-11 F3 and F4): a CORRECT current code returned that
// message in 0.0556 s with the hub firewalled off, and in 0.0299 s with this agent stopped —
// against ~1.0 s for a genuine unseal. No unseal was attempted in either case.
//
// 502 rather than 400: 4xx says "your request was bad", and the request was not bad — an
// upstream dependency failed. The status is the machine-readable half; the controller
// classifies on it and must never parse this sentence.
//
// ⚠ THE CODE WAS NOT USED. Nothing may be said about it — not that it was wrong, and not
// that it was right.
case errors.Is(err, escrow.ErrBundleFetch):
s.logger.Warn("local-api: offsite key recovery: the sealed bundle could not be FETCHED — the recovery code was never used", "vmid", vmid, "err", err)
writeErr(w, http.StatusBadGateway, "the sealed recovery bundle could not be fetched from the hub — the recovery code was NOT used and nothing was written")
case errors.Is(err, escrow.ErrNoEscrowBlob):
s.logger.Warn("local-api: offsite key recovery: the hub holds no sealed bundle for this host", "vmid", vmid)
writeErr(w, http.StatusNotFound, "the hub holds no sealed recovery bundle for this host — no escrow ceremony has run")
// ── R-311 (2026-08-12) — THE CODE IS RIGHT, JUST NOT FOR THE CURRENT PACKAGE. ─────────
//
// Placed ABOVE the default for the same reason ErrBundleFetch is: the default blames the
// customer, and this case is the one where the customer is provably not at fault. The code was
// used, it worked, and it opened a package the hub is deliberately keeping.
//
// 422 rather than 400: the request was well-formed AND the credential was valid — what could
// not be processed is the pairing of a correct code with the CURRENT package. A 400 would put
// it in the same bucket as a mistype, which is the whole defect. The status is the
// machine-readable half; the controller classifies on it and must never parse this sentence.
//
// The date travels in the body because it is the one fact that lets a customer recognise which
// code they are holding. No material, no code, no password — only when that package stopped
// being current, and whether it can yield a repository password at all.
case errors.Is(err, escrow.ErrCodeOpensRetained):
var ro *escrow.RetainedOpenedError
match := escrow.RetainedMatch{}
if errors.As(err, &ro) {
match = ro.Match
}
s.logger.Info("local-api: offsite key recovery: the code did NOT open the current package but DID open a RETAINED one — the customer is not at fault",
"vmid", vmid, "superseded_at", match.SupersededAt, "retained_has_restic_pw", match.HasResticPassword)
writeStatus(w, http.StatusUnprocessableEntity, false,
map[string]any{
"opens_retained": true,
"superseded_at": match.SupersededAt,
"retained_has_restic_pw": match.HasResticPassword,
},
"the recovery code is correct, but it belongs to an EARLIER sealed package (superseded "+match.SupersededAt+"), not the one currently held")
case errors.Is(err, escrow.ErrNoResticPassword):
s.logger.Warn("local-api: offsite key recovery: the bundle opened but predates the repository-password field", "vmid", vmid)
writeErr(w, http.StatusConflict, "the recovery code opened the bundle, but it carries NO offsite repository password (sealed before that field existed; it cannot be retro-fitted)")
default:
// The fail-closed WRONG-CODE case, and only it: the bundle was fetched and `age -d`
// refused it. Every other situation above has its own status. The agent log records the
// STEP, never the code.
s.logger.Warn("local-api: offsite key recovery: the fetched bundle did not open with the supplied recovery code", "vmid", vmid, "err", err)
writeErr(w, http.StatusBadRequest, "the recovery code did not open the sealed bundle — nothing was written")
}
return
}
sum := sha256.Sum256([]byte(strings.TrimSpace(pw)))
// §8.6's lesson, applied: say exactly WHAT was recovered and what was NOT, so nobody reading this
// concludes the wrong thing about the bundle's contents (which is how link 8 came to be missing).
s.logger.Info("local-api: offsite repository password RECOVERED from the sealed escrow — returning that field ONLY "+
"(the tunnel token, the PBS token and the WG key stay inside the agent and are not returned)",
"vmid", vmid, "restic_pw_sha256", hex.EncodeToString(sum[:]))
writeOK(w, map[string]any{
"restic_repo_password": pw,
"restic_pw_sha256": hex.EncodeToString(sum[:]),
})
}
@@ -0,0 +1,99 @@
package localapi
import (
"context"
"errors"
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/escrow"
)
// R-224 — THE STATUS IS THE DISCRIMINATOR, and this test asserts the CONSEQUENCE (what the HTTP
// boundary answers) rather than the mechanism (that the sentinel exists).
//
// The controller one trust tier down classifies on the STATUS and must never parse the sentence. So
// the contract this pins is: four distinguishable situations, four distinct statuses, and the
// wrong-code message reachable ONLY from a real refusal.
//
// Before R-224 the first and last rows both answered 400 with the same sentence — which is how
// CAMPAIGN-11 F3 told a customer holding a CORRECT code that it did not open their package.
type fakeRecoverer struct{ err error }
func (f fakeRecoverer) RecoverOffsiteRepoPassword(context.Context, string) (string, error) {
if f.err != nil {
return "", f.err
}
return "0123456789abcdef0123456789abcdef", nil
}
func TestRecoverOffsitePassword_EachSituationGetsItsOwnStatus(t *testing.T) {
cases := []struct {
name string
err error
wantStatus int
// mustNotSay guards the specific misattribution each status exists to prevent.
mustNotSay []string
}{
{
name: "fetch failed — the code was NEVER used",
err: errors.Join(escrow.ErrBundleFetch, errors.New("hub: transport error: no route to host")),
wantStatus: 502,
mustNotSay: []string{"did not open"},
},
{
name: "wrong code — the bundle WAS fetched and refused it",
err: errors.New("escrow: the recovery code did not unwrap the identity escrow"),
wantStatus: 400,
mustNotSay: []string{"could not be fetched"},
},
{
name: "the hub holds no bundle",
err: escrow.ErrNoEscrowBlob,
wantStatus: 404,
mustNotSay: []string{"did not open"},
},
{
name: "the bundle predates the repository-password field",
err: escrow.ErrNoResticPassword,
wantStatus: 409,
mustNotSay: []string{"could not be fetched"},
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
srv := newTestServerS(t, &fakeGuests{}, &fakeBackups{}, &fakeStore{}, nil)
srv.escrowRecovery = fakeRecoverer{err: tc.err}
w := do(t, srv.Handler(), "POST", "/escrow/recover-offsite-password", "A",
`{"vmid":8200,"recovery_code":"correct horse battery staple sedative anaconda wobbly kingdom placard yodel"}`)
if w.Code != tc.wantStatus {
t.Fatalf("status: got %d, want %d — body=%s", w.Code, tc.wantStatus, w.Body.String())
}
for _, phrase := range tc.mustNotSay {
if strings.Contains(w.Body.String(), phrase) {
t.Fatalf("the %d answer must not say %q — body=%s", tc.wantStatus, phrase, w.Body.String())
}
}
})
}
}
// The pair that matters most, stated as its own assertion so a regression cannot hide inside a table:
// a fetch failure and a wrong code must never answer with the SAME status. Collapsing them is the
// whole of R-224.
func TestRecoverOffsitePassword_FetchFailureAndWrongCodeDiffer(t *testing.T) {
status := func(err error) int {
srv := newTestServerS(t, &fakeGuests{}, &fakeBackups{}, &fakeStore{}, nil)
srv.escrowRecovery = fakeRecoverer{err: err}
return do(t, srv.Handler(), "POST", "/escrow/recover-offsite-password", "A",
`{"vmid":8200,"recovery_code":"correct horse battery staple sedative anaconda wobbly kingdom placard yodel"}`).Code
}
fetch := status(errors.Join(escrow.ErrBundleFetch, errors.New("no route to host")))
wrong := status(errors.New("escrow: the recovery code did not unwrap the identity escrow"))
// RED-PROOF: delete the ErrBundleFetch case from handleRecoverOffsitePassword → both become 400
// → this FAILS. That is the exact pre-R-224 code, and the exact defect CAMPAIGN-11 measured.
if fetch == wrong {
t.Fatalf("a failed fetch and a wrong code must not share a status (both %d)", fetch)
}
}
+117
View File
@@ -111,6 +111,14 @@ type HostMetricsProvider interface {
}
// Options configures a Server.
// EscrowRecoverer opens this host's sealed identity bundle with the customer recovery code and
// returns ONLY the offsite restic repository password (R-199 links 6-8). An interface so the
// localapi package needs no hub-client dependency and the route is testable without crypto.
// R is an argument and is never retained by any implementation.
type EscrowRecoverer interface {
RecoverOffsiteRepoPassword(ctx context.Context, recoveryCode string) (string, error)
}
type Options struct {
ListenAddr string // bridge IP:port
Cert tls.Certificate
@@ -177,6 +185,9 @@ type Options struct {
// StaleLock recovers a guest left with a stale vzdump lock by a reboot-during-backup (F2-b), run at
// startup by RecoverStaleLockedGuests. OPTIONAL — when nil, the recovery is a no-op.
StaleLock StaleLockController
// GuestsStateDir (R-523) is the agent's per-guest state dir holding <vmid>/bootstrap and the
// controller-parked marker. "" → /var/lib/felhom-agent/guests.
GuestsStateDir string
// ControllerSwapStateDir holds the per-guest swap state file (crash-safety). "" → /var/lib/felhom-agent.
ControllerSwapStateDir string
// Intent records drive enroll/eject intent for the self-heal watchdog (slice 10 P3). OPTIONAL —
@@ -210,6 +221,11 @@ type Options struct {
// GET /debug/logs. OPTIONAL — when nil the endpoint reports "not configured".
LogRing *applog.Ring
Logger *slog.Logger
// EscrowRecovery (R-199, v0.125.0) is the offsite-key recovery seam behind
// POST /escrow/recover-offsite-password. OPTIONAL — nil → that route reports "not configured"
// (503) instead of failing obscurely. Satisfied by escrow.OffsiteKeyRecoverer.
EscrowRecovery EscrowRecoverer
}
// defaultBackupCadence is the fallback /backup/due window when none is configured.
@@ -273,6 +289,11 @@ type Server struct {
netMountRoot string // the user-data namespace root for the network-mount role gate
smbCredsDir string // where SMB creds files are written (out-of-band, 0600)
escrowStagePath string // fork-4: 0600 staging file for the pushed restic repo password
// escrowRecovery (R-199, v0.125.0) assembles chain links 6-8: fetch this host's own sealed
// identity blob from the hub, unseal it with the customer's recovery code, return ONLY the
// offsite repository password. OPTIONAL — nil (no hub client configured) makes
// POST /escrow/recover-offsite-password answer 503 rather than pretending.
escrowRecovery EscrowRecoverer
intent IntentRecorder // slice 10 P3 (optional)
guestBinds *GuestBindStore // F9 startup bind re-assert record (optional)
formatJobs *FormatJobStore // F20-BUG3 detached-format job record (optional)
@@ -344,6 +365,12 @@ type Server struct {
swapMu sync.Mutex
swapInFlight map[int]bool
// R-523: the in-guest controller supervisor (controllersupervisor.go). guestExec is the same
// GuestExecutor the swap uses; guestsDir is the agent's per-guest state dir ("" → default).
guestExec GuestExecutor
guestsDir string
ctrlSup controllerSupervisor
// Network-storage verify job (SPIKE-nas-verify): the IN-MEMORY single slot + the seams the
// detached pipeline runs through (tests inject; production defaults set in NewServer).
netVerifyMu sync.Mutex
@@ -423,6 +450,7 @@ func NewServer(o Options) (*Server, error) {
netMountRoot: storage.NetworkMountRoot,
smbCredsDir: o.SmbCredsDir,
escrowStagePath: o.EscrowStagePath,
escrowRecovery: o.EscrowRecovery,
intent: o.Intent,
guestBinds: o.GuestBinds,
formatJobs: o.FormatJobs,
@@ -465,7 +493,9 @@ func NewServer(o Options) (*Server, error) {
s.statFile = func(path string) bool { _, err := os.Stat(path); return err == nil }
if o.ControllerSwap != nil {
s.swap = NewControllerSwapper(o.ControllerSwap, o.ControllerSwapStateDir, o.Logger)
s.guestExec = o.ControllerSwap
}
s.guestsDir = o.GuestsStateDir
return s, nil
}
@@ -518,6 +548,10 @@ func (s *Server) Handler() http.Handler {
mux.HandleFunc("POST /escrow/stage-secret", s.withGuest(s.handleStageEscrowSecret))
// fork-4 hygiene: wipe the staged secret once escrowed (controller calls this on confirm). Idempotent.
mux.HandleFunc("DELETE /escrow/stage-secret", s.withGuest(s.handleWipeStagedEscrowSecret))
// R-199 (v0.125.0): recover the offsite repository password from the hub-held sealed bundle,
// using the customer recovery code supplied in the body. Returns that ONE field. See
// escrow_recover.go for R's handling rules — they are the tightest in this package.
mux.HandleFunc("POST /escrow/recover-offsite-password", s.withGuest(s.handleRecoverOffsitePassword))
// Controller-driven escrow ceremony (v0.88.0): preflight checklist, the detached root ceremony
// job (fixed-argv sudo self-invocation), its status, and the ONE-SHOT in-memory R claim.
@@ -1061,6 +1095,10 @@ type BackupTierInfo struct {
Target string `json:"target"`
CadenceSeconds int64 `json:"cadence_seconds"`
Primary bool `json:"primary"`
// Storage (R-517/R-518, v0.131.0) says whether the tier's Proxmox storage exists on this host
// RIGHT NOW: "present" | "absent" | "unknown" (storage view unreadable). Additive — an older
// controller ignores it. "unknown" is never "absent": a probe failure must not skip a backup.
Storage string `json:"storage,omitempty"`
}
func (s *Server) handleBackupTiers(w http.ResponseWriter, r *http.Request, vmid int) {
@@ -1070,11 +1108,83 @@ func (s *Server) handleBackupTiers(w http.ResponseWriter, r *http.Request, vmid
Target: t.TargetID,
CadenceSeconds: int64(t.Cadence.Seconds()),
Primary: t.Primary,
Storage: s.storagePresence(r.Context(), t.TargetID),
})
}
writeOK(w, resp)
}
// storagePresence is the tri-state twin of targetStoragePresent (which must stay fail-OPEN for the
// backup path): "present", "absent", or "unknown" when the storage view cannot be read. Only a
// successful read that does not list the storage is "absent".
func (s *Server) storagePresence(ctx context.Context, target string) string {
if s.storage == nil || target == "" {
return StoragePresenceUnknown
}
targets, err := s.storage.Observe(ctx)
if err != nil {
s.logger.Warn("local-api: storage view unavailable for the tier presence report", "target", target, "err", err)
return StoragePresenceUnknown
}
for _, t := range targets {
if t.Name == target {
return StoragePresencePresent
}
}
return StoragePresenceAbsent
}
const (
StoragePresencePresent = "present"
StoragePresenceAbsent = "absent"
StoragePresenceUnknown = "unknown"
)
// TierBackupState (R-517, v0.131.0) is one tier's truth for the customer's backup page: the newest
// SUCCESSFUL backup and the last ATTEMPT, kept apart — so a failed attempt can never stand in for a
// result ("presence is not success").
type TierBackupState struct {
Target string `json:"target"`
Primary bool `json:"primary"`
Storage string `json:"storage"` // present | absent | unknown
// LastSuccess is the newest successful backup on this tier. From the in-memory record when there
// is one; otherwise from the tier's storage (after an agent restart the record is empty — the
// BIGNIGHT F2 page showed no backup at all), in which case only started_at is known and
// LastSuccessSource is "storage".
LastSuccess *hub.Backup `json:"last_success,omitempty"`
LastSuccessSource string `json:"last_success_source,omitempty"` // record | storage
LastAttempt *TierAttempt `json:"last_attempt,omitempty"`
}
// TierAttempt is the newest recorded attempt on a tier, successful or not.
type TierAttempt struct {
StartedAt string `json:"started_at"`
Success bool `json:"success"`
Error string `json:"error,omitempty"`
}
// tierBackupStates builds the per-tier view for one guest.
func (s *Server) tierBackupStates(ctx context.Context, vmid int) []TierBackupState {
out := make([]TierBackupState, 0, len(s.tiers))
for _, t := range s.tiers {
st := TierBackupState{Target: t.TargetID, Primary: t.Primary, Storage: s.storagePresence(ctx, t.TargetID)}
if b := s.pickLatestBackup(ctx, vmid, true, t.TargetID); b != nil {
st.LastSuccess, st.LastSuccessSource = b, "record"
} else if st.Storage != StoragePresenceAbsent {
if when, look := s.newestArchiveOn(ctx, t, vmid); look == archiveFound {
st.LastSuccess = &hub.Backup{TargetID: t.TargetID, VMID: vmid, Success: true,
StartedAt: when.UTC().Format(time.RFC3339)}
st.LastSuccessSource = "storage"
}
}
if a := s.pickLatestBackup(ctx, vmid, false, t.TargetID); a != nil {
st.LastAttempt = &TierAttempt{StartedAt: a.StartedAt, Success: a.Success, Error: a.Error}
}
out = append(out, st)
}
return out
}
// tierFromRequest resolves the `?target=` query parameter to a tier.
//
// THE COMPATIBILITY RULE (§4): NO target parameter → the PRIMARY tier, and the echoed target is
@@ -1121,6 +1231,10 @@ type BackupStatusResponse struct {
Backup *hub.Backup `json:"backup,omitempty"` // latest recorded backup for this guest
// Target (R-82) echoes the tier; empty + omitted when untargeted (pre-R-82 bytes).
Target string `json:"target,omitempty"`
// Tiers (R-517, v0.131.0) is the per-tier truth — newest success, last attempt, storage
// presence. Served on the UNTARGETED request only; additive, so an older controller reads the
// response exactly as before.
Tiers []TierBackupState `json:"tiers,omitempty"`
}
func (s *Server) handleBackupStatus(w http.ResponseWriter, r *http.Request, vmid int) {
@@ -1132,6 +1246,9 @@ func (s *Server) handleBackupStatus(w http.ResponseWriter, r *http.Request, vmid
// across ANY target (echo == "" → pickLatestBackup's match-any path).
resp := BackupStatusResponse{VMID: vmid, Phase: PhaseIdle, Target: echo,
Backup: s.pickLatestBackup(r.Context(), vmid, false, echo)}
if echo == "" {
resp.Tiers = s.tierBackupStates(r.Context(), vmid)
}
if job, ok := s.jobSnapshot(backupJobKey{vmid: vmid, target: tier.TargetID}); ok {
resp.Phase = job.Phase
resp.JobID = job.JobID
+14 -2
View File
@@ -10,6 +10,8 @@ import (
"net/url"
"strings"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/httpx"
)
// Client is the PBS-API client for ONE PBS server. Construct with NewClient. It is pure (no
@@ -31,6 +33,11 @@ type Config struct {
Secret string // token secret (from <id>.pw)
Namespace string // PBS namespace (from storage.cfg `namespace`); "" = root. S4 per-customer tenancy.
Timeout time.Duration
// IdleConnTimeout bounds how long this client's idle keep-alive connections are retained.
// Zero means httpx.DefaultIdleConnTimeout (90s) — it does NOT mean "no timeout", which is the
// R-344 defect. Production leaves it unset; only tests set it, to avoid a 90-second wait.
IdleConnTimeout time.Duration
}
// NewClient builds a fingerprint-pinned, token-authed PBS client.
@@ -55,8 +62,13 @@ func NewClient(cfg Config) (*Client, error) {
authHeader: "PBSAPIToken=" + cfg.TokenID + ":" + cfg.Secret,
namespace: cfg.Namespace,
http: &http.Client{
Timeout: timeout,
Transport: &http.Transport{TLSClientConfig: tlsCfg},
Timeout: timeout,
// R-344: this transport MUST come from httpx. pbsTargetsFromPVE (cmd/felhom-agent/
// main.go) builds a fresh Client every cycle and drops the previous one, so a
// transport with no idle timeout strands one connection per cycle, forever, on both
// sides. That leaked 388 sockets onto ep0 in 46 hours. Pinned by
// TestAbandonedClientsReleaseTheirConnections — do not inline an http.Transport here.
Transport: httpx.NewTransport(tlsCfg, cfg.IdleConnTimeout),
},
}, nil
}
+173
View File
@@ -0,0 +1,173 @@
package pbs
import (
"context"
"net"
"net/http"
"net/http/httptest"
"strings"
"sync"
"testing"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/httpx"
)
// connCounter is the SERVER-side observer. It counts what the server actually holds, which is the
// only thing that answers the question this file exists for: a client that believes it closed a
// connection, and a server still holding the socket, is precisely the R-344 shape. Asserting on
// anything client-side would be asserting the mechanism instead of the consequence.
type connCounter struct {
mu sync.Mutex
open int
total int // every connection ever accepted — how many times the client DIALLED
}
func (c *connCounter) hook(_ net.Conn, s http.ConnState) {
c.mu.Lock()
defer c.mu.Unlock()
switch s {
case http.StateNew:
c.open++
c.total++
case http.StateClosed, http.StateHijacked:
c.open--
}
}
func (c *connCounter) counts() (open, total int) {
c.mu.Lock()
defer c.mu.Unlock()
return c.open, c.total
}
// waitForOpen polls until the server holds want connections, or fails naming what it still holds.
func (c *connCounter) waitForOpen(t *testing.T, want int, within time.Duration, what string) {
t.Helper()
deadline := time.Now().Add(within)
for {
open, total := c.counts()
if open == want {
return
}
if time.Now().After(deadline) {
t.Fatalf("%s: after %s the server still holds %d open connection(s), want %d (%d dialled in total)",
what, within, open, want, total)
}
time.Sleep(5 * time.Millisecond)
}
}
// newCountingPBSServer is newPBSTestServer plus a ConnState hook. Kept separate rather than
// changing the shared helper, so the existing tests are untouched by this file.
func newCountingPBSServer(t *testing.T, fn http.HandlerFunc) (*httptest.Server, string, *connCounter) {
t.Helper()
cc := &connCounter{}
ts := httptest.NewUnstartedServer(fn)
ts.Config.ConnState = cc.hook
ts.StartTLS()
t.Cleanup(ts.Close)
return ts, fingerprintOf(ts), cc
}
// TestAbandonedClientsReleaseTheirConnections is Scenario A, and it is the load-bearing test for
// R-344.
//
// It models what pbsTargetsFromPVE actually does — build a client, use it once, drop it on the
// floor without closing anything — and asserts the CONSEQUENCE on the server: the connections go
// away. Before the fix every one of these stayed established forever on both sides; 388 of them
// accumulated on ep0 in 46 hours.
//
// Deliberately NOT asserted: that err == nil, or that IdleConnTimeout holds some value. Both were
// true of the leaking code.
func TestAbandonedClientsReleaseTheirConnections(t *testing.T) {
ts, fp, cc := newCountingPBSServer(t, func(w http.ResponseWriter, _ *http.Request) {
w.Write([]byte(`{"data":[]}`))
})
host, port := hostPort(t, ts.URL)
const cycles = 5
for i := 0; i < cycles; i++ {
// One fresh client per "cycle", exactly as pbsTargetsFromPVE builds one per collect.
c, err := NewClient(Config{
Server: host, Port: port, Fingerprint: fp, TokenID: "u@pbs!t", Secret: "s",
IdleConnTimeout: 50 * time.Millisecond, // production uses the 90s default
})
if err != nil {
t.Fatal(err)
}
if _, err := c.Snapshots(context.Background(), "ds"); err != nil {
t.Fatalf("cycle %d: %v", i, err)
}
_ = c // dropped here — nothing closes it, nothing can
}
if _, total := cc.counts(); total != cycles {
t.Fatalf("setup is not modelling the leak: want %d separate dials (one per abandoned client), got %d", cycles, total)
}
cc.waitForOpen(t, 0, 5*time.Second, "abandoned pbs.Clients")
}
// TestAbandonedClientsReleaseTheirConnections_ProductionDefaultIsUsable pins the value that ships.
//
// The field being settable is exactly how it could silently become zero again — and zero used to
// mean "never expire". This asserts the production path (Config leaving it unset) lands on the
// standard-library default, so the leak cannot return through an unset field.
func TestPBSClient_UnsetIdleTimeoutUsesTheDefault(t *testing.T) {
for _, tc := range []struct {
name string
cfg time.Duration
want time.Duration
}{
{"unset — the production path", 0, httpx.DefaultIdleConnTimeout},
{"explicit zero is NOT no-timeout", 0, httpx.DefaultIdleConnTimeout},
{"negative is NOT no-timeout", -time.Second, httpx.DefaultIdleConnTimeout},
{"an explicit value is honoured", 3 * time.Second, 3 * time.Second},
} {
t.Run(tc.name, func(t *testing.T) {
c, err := NewClient(Config{
Server: "pbs.example", Fingerprint: strings.Repeat("ab", 32),
TokenID: "u@pbs!t", Secret: "s", IdleConnTimeout: tc.cfg,
})
if err != nil {
t.Fatal(err)
}
tr, ok := c.http.Transport.(*http.Transport)
if !ok {
t.Fatalf("transport is %T, not *http.Transport — the httpx wiring was replaced", c.http.Transport)
}
if tr.IdleConnTimeout != tc.want {
t.Fatalf("IdleConnTimeout = %v, want %v (zero would mean connections are retained FOREVER — that is R-344)", tr.IdleConnTimeout, tc.want)
}
})
}
}
// TestPBSClient_KeepAliveStillReuses is Scenario C, and it is the guard against a "fix" that is
// worse than the bug.
//
// Disabling keep-alive entirely would also make the leak go away — by dialling a fresh connection
// for every single request, which on a box polling ~40,000 times a day is strictly worse than what
// we started with. The fix must retire IDLE connections without stopping reuse.
func TestPBSClient_KeepAliveStillReuses(t *testing.T) {
ts, fp, cc := newCountingPBSServer(t, func(w http.ResponseWriter, _ *http.Request) {
w.Write([]byte(`{"data":[]}`))
})
host, port := hostPort(t, ts.URL)
c, err := NewClient(Config{
Server: host, Port: port, Fingerprint: fp, TokenID: "u@pbs!t", Secret: "s",
IdleConnTimeout: 30 * time.Second, // long enough that reuse is what is being measured
})
if err != nil {
t.Fatal(err)
}
for i := 0; i < 3; i++ {
if _, err := c.Snapshots(context.Background(), "ds"); err != nil {
t.Fatalf("request %d: %v", i, err)
}
}
if _, total := cc.counts(); total != 1 {
t.Fatalf("one client made 3 sequential requests over %d connection(s), want 1 — keep-alive reuse is broken, which would make the poll load WORSE than the leak", total)
}
}
+30 -1
View File
@@ -283,7 +283,36 @@ func (m *Manager) Apply(ctx context.Context, fetched bool, block *hub.WirePBSDR)
h := descriptorHash(block)
cf := m.loadConsumedFailed()
if mk := m.loadMarker(); mk != nil && mk.Hash == h && (cf == nil || cf.Hash != h) {
m.setStatus(&hub.PBSDRStatus{State: mk.State, StorageID: block.StorageID, Namespace: block.Namespace, AppliedAt: mk.AppliedAt})
// R-221: RE-ASSERT THE SEED, DO NOT REMEMBER IT. The marker records that this descriptor
// converged; it says nothing about whether the file the seed writes still exists.
//
// The two live in different places and die at different times. The marker is host-side
// (`<agent-state>/pbsdr/`, markerPath above); the seed's target is `agent.json`, and the
// installer's `step_agent_config` renders that file from `base = {}` unless an explicit
// `--preserve-from` is passed — it NEVER writes an `escrow` section — then replaces it with
// O_TRUNC (felhom-host-install.sh:2396, :2449, :2579; the flag is :1246, defaulting empty at
// :256). So a rebuild leaves the marker and takes the seed, the hash still matches, this
// branch returns, and `escrow.pbs_storage_id` is never written again. The customer then
// cannot run the escrow ceremony AT ALL: handleEscrowPreflight fails the `pbs_storage_id`
// row and the wizard refuses, with no way forward from inside the product.
//
// A rebuild is only the case that was measured. The same hole opens for a hand-edited or
// restored config, which is the honest reason this is a seam fix rather than an installer
// fix — the seed must be a thing the loop asserts, not a thing it did once.
//
// COST: this runs on the converged path, i.e. every tick (60 s) forever. It is one small
// file read plus a JSON parse — no exec, no network, no Proxmox call — and seedEscrowStorageID
// returns early once the value matches. That is the whole reason it is affordable here.
//
// IT MUST NEVER UN-CONVERGE THE BOX: a failure is a Warn plus a message on the published
// status, exactly as finishConverged does it. No marker write, no state change, no retry
// storm — the early return below still happens either way.
msg := ""
if err := m.seedEscrowStorageID(block.StorageID); err != nil {
msg = "escrow.pbs_storage_id seed failed: " + err.Error() + " (set it manually before the ceremony)"
m.logger.Warn("pbsdr: " + msg)
}
m.setStatus(&hub.PBSDRStatus{State: mk.State, StorageID: block.StorageID, Namespace: block.Namespace, AppliedAt: mk.AppliedAt, Message: msg})
return // idempotent: this exact descriptor already converged
}
+232
View File
@@ -0,0 +1,232 @@
package pbsdr
import (
"context"
"encoding/json"
"go/ast"
"go/parser"
"go/printer"
"go/token"
"io"
"os"
"path/filepath"
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
)
// R-221 — the escrow seed must be ASSERTED on every converged tick, not remembered.
//
// These tests drive the real Apply() with a real temp-dir agent.json and a call-recording runner.
// Calling seedEscrowStorageID directly would prove nothing: the defect IS the early return in
// Apply, and a test that steps around it cannot see it.
// convergedMarker writes a marker whose hash matches the block, i.e. puts the manager on exactly
// the idempotent path where the seed used to be skipped.
func convergedMarker(t *testing.T, m *Manager, block *hub.WirePBSDR, state string) {
t.Helper()
if err := m.writeState(m.markerPath(), marker{
Hash: descriptorHash(block), State: state, AppliedAt: "2026-08-08T00:00:00Z",
}); err != nil {
t.Fatalf("write marker: %v", err)
}
}
func escrowStorageID(t *testing.T, cfgPath string) string {
t.Helper()
raw, err := os.ReadFile(cfgPath)
if err != nil {
t.Fatalf("read config: %v", err)
}
var doc struct {
Escrow struct {
PBSStorageID string `json:"pbs_storage_id"`
} `json:"escrow"`
}
if err := json.Unmarshal(raw, &doc); err != nil {
t.Fatalf("parse config: %v", err)
}
return doc.Escrow.PBSStorageID
}
// drBlock mirrors the existing valid fixture (manager_test.go descriptor()) so that validate()
// passes and Apply actually reaches the marker check — the branch these tests are about.
func drBlock(storageID string) *hub.WirePBSDR {
return &hub.WirePBSDR{
Enabled: true, StorageID: storageID, PBSTunnelIP: "10.77.0.1",
Datastore: "felhom-offsite", Namespace: "peti", TokenID: "felhom@pbs!peti",
Fingerprint: testFP,
}
}
// SCENARIO A — the seed is re-asserted on a converged box, and nothing else happens.
//
// THIS IS THE TEST THAT MATTERS. It must fail against the pre-R-221 tree; if it passes there, it is
// not testing the defect and that is the finding.
func TestSeedReasserted_OnConvergedTick_WithZeroProxmoxCalls(t *testing.T) {
r := &fakeRunner{}
st := &fakeStorage{found: true, active: []bool{true}}
c := &fakeConsumer{}
m, cfgPath := newTestManager(t, r, st, c)
block := drBlock("felhom-pbs-dr")
convergedMarker(t, m, block, "applied")
// the shape a rebuild leaves behind: escrow section present, pbs_storage_id GONE
if got := escrowStorageID(t, cfgPath); got != "" {
t.Fatalf("precondition: config already carries a storage id %q", got)
}
m.Apply(context.Background(), true, block)
if got := escrowStorageID(t, cfgPath); got != "felhom-pbs-dr" {
t.Errorf("the converged tick did not re-assert the seed: escrow.pbs_storage_id = %q, want %q.\n"+
"This is R-221: the marker survives a rebuild, the descriptor hash still matches, the early "+
"return fires and the seed never runs into the config that no longer has it — so the customer "+
"cannot run the escrow ceremony at all.", got, "felhom-pbs-dr")
}
// ...and the idempotent path is STILL idempotent. This assertion is not decorative: without it
// a "fix" that simply deletes the early return would pass the line above.
if calls := r.recorded(); len(calls) != 0 {
t.Errorf("a converged tick must execute ZERO Proxmox commands; got %d: %+v", len(calls), calls)
}
}
// SCENARIO B — an operator's own different value survives, and the warning names both.
func TestSeedReasserted_NeverClobbersAnOperatorValue(t *testing.T) {
r := &fakeRunner{}
m, cfgPath := newTestManager(t, r, &fakeStorage{found: true, active: []bool{true}}, &fakeConsumer{})
if err := os.WriteFile(cfgPath, []byte(
`{"log_level":"info","escrow":{"posture":"zero_knowledge","pbs_storage_id":"operator-chosen"},`+
`"custom_unknown":{"keep":1}}`), 0o600); err != nil {
t.Fatal(err)
}
block := drBlock("hub-chosen")
convergedMarker(t, m, block, "applied")
m.Apply(context.Background(), true, block)
if got := escrowStorageID(t, cfgPath); got != "operator-chosen" {
t.Errorf("a value a person put there was overwritten by the descriptor: got %q, want %q", got, "operator-chosen")
}
// unknown keys must still round-trip
raw, _ := os.ReadFile(cfgPath)
if !strings.Contains(string(raw), "custom_unknown") {
t.Error("an unknown config key was dropped by the re-assert")
}
}
// SCENARIO C — the ceremony preflight's live read sees the re-asserted value with NO restart.
//
// The preflight itself lives in internal/localapi and reads the file through config.Load; what this
// asserts is the half that belongs to this package: after a converged tick, THE FILE ON DISK carries
// the id, so any live re-read is green. The daemon is never restarted in this test because it is
// never started — which is the point.
func TestSeedReasserted_IsVisibleOnDiskImmediately(t *testing.T) {
r := &fakeRunner{}
m, cfgPath := newTestManager(t, r, &fakeStorage{found: true, active: []bool{true}}, &fakeConsumer{})
block := drBlock("felhom-pbs-dr")
convergedMarker(t, m, block, "applied")
// preflight's predicate BEFORE: storageID == "" → the row is NOT OK
if escrowStorageID(t, cfgPath) != "" {
t.Fatal("precondition")
}
m.Apply(context.Background(), true, block)
// preflight's predicate AFTER, from the same file the ceremony subprocess loads
if id := escrowStorageID(t, cfgPath); id == "" {
t.Error("the preflight row would still be NOT OK after a converged tick")
}
}
// A seed failure must NEVER un-converge the box: no marker rewrite, no state change, and the status
// still reports the marker's converged state — with the failure surfaced as a message.
func TestSeedReassertFailure_DoesNotUnconverge(t *testing.T) {
r := &fakeRunner{}
m, cfgPath := newTestManager(t, r, &fakeStorage{found: true, active: []bool{true}}, &fakeConsumer{})
block := drBlock("felhom-pbs-dr")
convergedMarker(t, m, block, "applied")
markerBefore, err := os.ReadFile(m.markerPath())
if err != nil {
t.Fatal(err)
}
// make the seed fail in a way it cannot recover from: unparseable config
if err := os.WriteFile(cfgPath, []byte(`{ this is not json`), 0o600); err != nil {
t.Fatal(err)
}
m.Apply(context.Background(), true, block)
after, err := os.ReadFile(m.markerPath())
if err != nil {
t.Fatalf("the marker was removed by a seed failure: %v", err)
}
if string(after) != string(markerBefore) {
t.Error("a seed failure rewrote the convergence marker — it must not touch state")
}
if calls := r.recorded(); len(calls) != 0 {
t.Errorf("a seed failure must not trigger Proxmox work; got %+v", calls)
}
st := m.Status()
if st == nil || st.State != "applied" {
t.Errorf("a seed failure must leave the box converged; status = %+v", st)
}
if st != nil && !strings.Contains(st.Message, "seed failed") {
t.Errorf("a seed failure must be surfaced on the status, got message %q", st.Message)
}
}
// SEAM WIRING — production must construct the manager with the live config path, or the whole seed
// leg is inert. Three shipped defects in this project were fully green while their seam was never
// wired, so this walks main.go's AST for the actual call rather than grepping: a commented-out call
// satisfies strings.Contains, and an AST walk cannot see a comment.
func TestProductionWiring_NewManagerGetsTheLiveConfigPath(t *testing.T) {
path := filepath.Join("..", "..", "cmd", "felhom-agent", "main.go")
fset := token.NewFileSet()
f, err := parser.ParseFile(fset, path, nil, 0) // comments not even collected
if err != nil {
t.Fatalf("parse main.go: %v", err)
}
found := false
ast.Inspect(f, func(n ast.Node) bool {
call, ok := n.(*ast.CallExpr)
if !ok {
return true
}
sel, ok := call.Fun.(*ast.SelectorExpr)
if !ok || sel.Sel.Name != "NewManager" {
return true
}
if pkg, ok := sel.X.(*ast.Ident); !ok || pkg.Name != "pbsdr" {
return true
}
// signature: (runner, px, hubc, stateDir, secretDir, configPath, logger)
if len(call.Args) < 6 {
t.Errorf("pbsdr.NewManager called with %d args, expected >= 6", len(call.Args))
return false
}
var buf strings.Builder
if err := printNode(&buf, fset, call.Args[5]); err != nil {
t.Fatalf("print arg: %v", err)
}
got := buf.String()
if !strings.Contains(got, "SourcePath") {
t.Errorf("pbsdr.NewManager's configPath argument is %q, which is not the live config path.\n"+
"With an empty or wrong path seedEscrowStorageID returns nil immediately and the entire "+
"R-221 fix is inert while every test above still passes.", got)
}
found = true
return false
})
if !found {
t.Error("no pbsdr.NewManager call found in main.go — the manager is not constructed in production")
}
}
func printNode(w io.Writer, fset *token.FileSet, n ast.Node) error {
return printer.Fprint(w, fset, n)
}
+6 -2
View File
@@ -10,6 +10,8 @@ import (
"net/url"
"strings"
"time"
"gitea.dooplex.hu/admin/felhom-agent/internal/httpx"
)
// doer is the minimal HTTP surface the client needs; *http.Client satisfies it.
@@ -65,8 +67,10 @@ func NewClient(cfg Config) (*Client, error) {
timeout = 30 * time.Second
}
hc := &http.Client{
Timeout: timeout,
Transport: &http.Transport{TLSClientConfig: tlsCfg},
Timeout: timeout,
// R-344, consistency only: built ONCE per process, so it never accumulated and contributed
// nothing to the ep0 leak. Same missing default, corrected for the same reason.
Transport: httpx.NewTransport(tlsCfg, 0),
}
return &Client{
base: strings.TrimRight(cfg.Endpoint, "/") + "/api2/json",
+1 -1
View File
@@ -52,7 +52,7 @@ func newDREngine(t *testing.T, api GuestAPI) (*Engine, *fakeRunner, string, *Que
t.Cleanup(q.Close)
fr := &fakeRunner{}
sd := t.TempDir()
e := NewEngine(EngineOptions{API: api, Queue: q, Journal: j, Provider: EmptyProvider{}, HostRunner: fr, StateDir: sd})
e := NewEngine(EngineOptions{API: api, Queue: q, Journal: j, Provider: EmptyProvider{}, HostRunner: fr, StateDir: sd, RestoreSpace: roomySpace{}})
return e, fr, sd, q
}
+30
View File
@@ -44,6 +44,17 @@ type Engine struct {
opSeq uint64 // atomic; makes each op id unique per attempt
// restoreSpace + spacePolicy are the restore-test's space preflight (R-672). nil space REFUSES
// every restore-test (fail-closed) — see restoretest_space.go.
restoreSpace RestoreSpace
spacePolicy SpacePolicy
// scratchMu guards activeScratch (the vmids a running restore-test owns — the retry timer never
// touches those) and teardownTries (failed timer retries per journal op, R-672 rule 3).
scratchMu sync.Mutex
activeScratch map[int]bool
teardownTries map[string]int
// lastRes records the most recent successful Reconcile Result (v0.90.0, R-28 fast-tick source).
// The fast-tick reads it to decide convergence: actionable drift is Planned − Pending > 0 (a
// destructive pending_signature refusal is EXPECTED state, not drift to hammer on). lastOK is
@@ -86,6 +97,10 @@ type EngineOptions struct {
HostRunner proxmox.Runner
// StateDir is the agent state dir ("" → /var/lib/felhom-agent); only the 4d swap reads it.
StateDir string
// RestoreSpace is the restore-test's space preflight (R-672). nil → every restore-test is REFUSED
// with its reason (fail-closed). SpacePolicy zero → DefaultSpacePolicy.
RestoreSpace RestoreSpace
SpacePolicy SpacePolicy
}
// NewEngine builds an Engine. The Queue is shared (the single §10 choke point); the
@@ -124,9 +139,24 @@ func NewEngine(opts EngineOptions) *Engine {
logger: logger,
hostRun: opts.HostRunner,
stateDir: stateDir,
restoreSpace: opts.RestoreSpace,
spacePolicy: policyOrDefault(opts.SpacePolicy),
activeScratch: map[int]bool{},
teardownTries: map[string]int{},
}
}
func policyOrDefault(p SpacePolicy) SpacePolicy {
if p.Factor < 1 {
p.Factor = DefaultSpacePolicy.Factor
}
if p.ReserveBytes <= 0 {
p.ReserveBytes = DefaultSpacePolicy.ReserveBytes
}
return p
}
// Result summarizes one Reconcile pass.
type Result struct {
Planned int
+1 -1
View File
@@ -213,7 +213,7 @@ func newEngine(t *testing.T, api GuestAPI, provider DesiredProvider) (*Engine, *
t.Cleanup(func() { j.Close() })
q := NewQueue()
t.Cleanup(q.Close)
e := NewEngine(EngineOptions{API: api, Queue: q, Journal: j, Provider: provider})
e := NewEngine(EngineOptions{API: api, Queue: q, Journal: j, Provider: provider, RestoreSpace: roomySpace{}})
return e, j, q
}
+40 -11
View File
@@ -50,10 +50,18 @@ type RestoreTestResult struct {
ScratchVMID int
Pass bool
Verified string // "boot+running" this slice
Skipped bool // no free scratch VMID in band → test not run
Err error
StartedAt time.Time
Duration time.Duration
Skipped bool // test not run: no free scratch VMID in band, or the space preflight refused (R-672)
// SkipReason is set when the SPACE PREFLIGHT refused (R-672): the test did not run, and this is
// reported to the hub as the test's result (pass=false), never as a pass. Empty for a band skip.
SkipReason string
// TargetStorage is where the restore went (rule 2 may move it off the tested guest's pool);
// RequiredBytes/AvailBytes are rule 1's figures.
TargetStorage string
RequiredBytes int64
AvailBytes int64
Err error
StartedAt time.Time
Duration time.Duration
// StartWarnings holds the warning line(s) the guest-start task emitted (e.g. the
// systemd-nesting advisory). Populated only when the start exited "WARNINGS: N";
// always surfaced, NEVER used to decide pass/fail (the verdict is liveness — waitRunning).
@@ -136,6 +144,29 @@ func (e *Engine) RunRestoreTest(ctx context.Context, spec RestoreTestSpec) Resto
return res
}
// R-672: the space preflight, BEFORE anything is journaled or created.
rawCfg, err := e.api.ExtractArchiveConfig(ctx, spec.Archive)
if err != nil {
res.Err = fmt.Errorf("reconcile: restore-test extract archive config: %w", err)
return res
}
v := PreflightRestoreSpace(ctx, e.restoreSpace, e.spacePolicy, spec.Archive, rawCfg, spec.RestoreStorage)
res.TargetStorage, res.RequiredBytes, res.AvailBytes = v.Storage, v.Required, v.Avail
if !v.OK {
res.Skipped = true
res.SkipReason = "skipped: " + v.Reason
e.logger.Warn("restore-test SKIPPED by the space preflight (R-672) — nothing was created",
"archive", spec.Archive, "storage", v.Storage, "required_bytes", v.Required, "avail_bytes", v.Avail, "reason", v.Reason)
res.Duration = time.Since(now)
return res
}
if v.Storage != spec.RestoreStorage {
e.logger.Info("restore-test: restoring OFF the tested guest's own pool (R-672 rule 2)",
"configured", spec.RestoreStorage, "avoided", v.Avoided, "target", v.Storage)
}
e.logger.Info("restore-test: space preflight passed", "storage", v.Storage, "required_bytes", v.Required, "avail_bytes", v.Avail)
spec.RestoreStorage = v.Storage
lxc, err := e.api.ListLXC(ctx)
if err != nil {
res.Err = fmt.Errorf("reconcile: restore-test list guests: %w", err)
@@ -164,7 +195,9 @@ func (e *Engine) RunRestoreTest(ctx context.Context, spec RestoreTestSpec) Resto
// Serialize on the scratch VMID's lane (inherits §10), and capture the result.
var vmidOccupied bool
ch := e.queue.Submit(vmid, func() error {
vmidOccupied = e.runScratchTest(ctx, vmid, spec, &res)
e.markScratch(vmid, true)
defer e.markScratch(vmid, false)
vmidOccupied = e.runScratchTest(ctx, vmid, spec, rawCfg, &res)
return res.Err
})
<-ch
@@ -182,7 +215,7 @@ func (e *Engine) RunRestoreTest(ctx context.Context, spec RestoreTestSpec) Resto
// runScratchTest is the journaled body (runs on vmid's queue lane). The occupied return is true
// ONLY when PVE synchronously refused the restore because the vmid already holds a guest (one
// the pool-blind band scan couldn't see) — the caller then advances to the next band vmid (F2).
func (e *Engine) runScratchTest(ctx context.Context, vmid int, spec RestoreTestSpec, res *RestoreTestResult) (occupied bool) {
func (e *Engine) runScratchTest(ctx context.Context, vmid int, spec RestoreTestSpec, rawCfg string, res *RestoreTestResult) (occupied bool) {
base := JournalEntry{OpID: e.scratchOpID(vmid), VMID: vmid, Kind: scratchKind, Scratch: true}
// OWN the scratch guest's cleanup BEFORE any mutation. From here, a crash is recoverable.
@@ -215,11 +248,7 @@ func (e *Engine) runScratchTest(ctx context.Context, vmid int, spec RestoreTestS
// genuinely EXTRACTED — full fidelity; the added runtime IS the verification), the two
// structural binds → throwaway stand-ins. An unreadable archive config or an unknown
// topology REFUSES up front — never restore a partial guest to "verify" it.
rawCfg, err := e.api.ExtractArchiveConfig(ctx, spec.Archive)
if err != nil {
res.Err = fmt.Errorf("reconcile: restore-test extract archive config: %w", err)
return false
}
// The archive's config was read ONCE, by the space preflight (R-672), and is passed in.
mountOverrides, err := drRestoreOverrides(rawCfg, spec.RestoreStorage)
if err != nil {
res.Err = fmt.Errorf("reconcile: restore-test: %w", err)
+94
View File
@@ -0,0 +1,94 @@
package reconcile
import "context"
// ── A failed scratch teardown is retried on a TIMER, not only at agent start (R-672 rule 3) ────────
//
// MEASURED 2026-09-24 on demo-hp: the scheduled restore-test's teardown failed (`lvremove … contains a
// filesystem in use`, a transient hold) and logged "left for Recover" — and Recover runs ONLY at agent
// start, so the 22 GiB scratch guest sat in the full pool for 2.5 hours until an agent restart. The
// timer calls RetryScratchTeardown every 10 minutes: the SAME resolution as Recover (recoverScratch —
// the gate's benign scratch destroy, idempotent when the guest is already gone), restricted to Scratch
// entries that carry a launch-proof UPID and that no running restore-test owns. After
// MaxTeardownTries failed attempts for one entry the operator is told (the caller reports it); the
// timer keeps trying.
// MaxTeardownTries is how many failed timer retries of one scratch entry happen before the operator
// is told.
const MaxTeardownTries = 3
// ScratchRetryResult summarizes one timer pass.
type ScratchRetryResult struct {
Examined int
Destroyed int
Clean int // already gone
Failed int
// GaveUp lists the scratch vmids whose failed tries reached MaxTeardownTries IN THIS PASS — each
// is reported exactly once (the caller tells the operator).
GaveUp []int
}
func (e *Engine) markScratch(vmid int, active bool) {
e.scratchMu.Lock()
defer e.scratchMu.Unlock()
if active {
e.activeScratch[vmid] = true
} else {
delete(e.activeScratch, vmid)
}
}
func (e *Engine) scratchActive(vmid int) bool {
e.scratchMu.Lock()
defer e.scratchMu.Unlock()
return e.activeScratch[vmid]
}
// RetryScratchTeardown is the timer's pass. It never touches a non-Scratch entry (unlike Recover,
// which also resolves generic in-flight operations and must therefore run only at start), never an
// entry without a launch-proof UPID (nothing was created), and never a vmid a running test owns.
func (e *Engine) RetryScratchTeardown(ctx context.Context) ScratchRetryResult {
var out ScratchRetryResult
if e.journal == nil {
return out
}
for _, entry := range e.journal.InFlight() {
if !entry.Scratch || entry.UPID == "" || e.scratchActive(entry.VMID) {
continue
}
out.Examined++
var r RecoverResult
e.recoverScratch(ctx, entry, &r)
switch {
case r.ScratchDestroyed > 0:
out.Destroyed++
e.forgetTries(entry.OpID)
case r.ScratchClean > 0:
out.Clean++
e.forgetTries(entry.OpID)
default:
out.Failed++
n := e.addTry(entry.OpID)
e.logger.Warn("restore-test: scratch teardown retry failed (timer)", "vmid", entry.VMID, "op_id", entry.OpID, "try", n)
if n == MaxTeardownTries {
out.GaveUp = append(out.GaveUp, entry.VMID)
e.logger.Error("restore-test: scratch guest still NOT torn down after repeated retries — telling the operator",
"vmid", entry.VMID, "tries", n)
}
}
}
return out
}
func (e *Engine) addTry(op string) int {
e.scratchMu.Lock()
defer e.scratchMu.Unlock()
e.teardownTries[op]++
return e.teardownTries[op]
}
func (e *Engine) forgetTries(op string) {
e.scratchMu.Lock()
defer e.scratchMu.Unlock()
delete(e.teardownTries, op)
}
@@ -0,0 +1,73 @@
package reconcile
import (
"context"
"errors"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-672 rule 3 (v0.133.0): a failed scratch teardown is retried on a TIMER. The consequence asserted:
// the leaked scratch guest is destroyed by a timer pass (not only by a restart's Recover), the operator
// is told exactly once after MaxTeardownTries failures, and a vmid a running test owns is never touched.
// leakScratch runs a restore-test whose teardown fails, leaving scratch 990000 in-flight — the
// 2026-09-24 shape ("lvremove … contains a filesystem in use").
func leakScratch(t *testing.T) (*Engine, *fakeAPI, *Journal) {
t.Helper()
api := &fakeAPI{cfg: map[int]proxmox.GuestConfig{990000: scratchCfg()}, restoreUPID: "UPID:r", destroyErr: errors.New("lvremove: contains a filesystem in use")}
e, j := spaceEngine(t, api, roomySpace{})
e.RunRestoreTest(context.Background(), RestoreTestSpec{Archive: "local:backup/x.tar.zst", RestoreStorage: "local-lvm", ScratchMin: 990000, ScratchMax: 990009})
if len(j.InFlight()) != 1 {
t.Fatalf("setup: want the scratch left in-flight after a failed teardown, got %+v", j.InFlight())
}
api.lxc = []proxmox.Guest{{VMID: 990000}}
return e, api, j
}
// COMPANION RED-PROOF (REPORT): make RetryScratchTeardown return without touching the journal (the
// v0.132.0 shape — only Recover at start resolved a leak) → "the leaked scratch was not destroyed by
// the timer".
func TestRetry_TheTimerDestroysALeakedScratch(t *testing.T) {
e, api, j := leakScratch(t)
api.destroyErr = nil // the transient hold is gone
before := len(api.destroys)
r := e.RetryScratchTeardown(context.Background())
if r.Destroyed != 1 || len(api.destroys) != before+1 || api.destroys[len(api.destroys)-1] != 990000 {
t.Fatalf("the leaked scratch was not destroyed by the timer: result=%+v destroys=%v", r, api.destroys)
}
if len(j.InFlight()) != 0 {
t.Fatalf("the entry is still in flight after a successful retry: %+v", j.InFlight())
}
if r2 := e.RetryScratchTeardown(context.Background()); r2.Examined != 0 {
t.Fatalf("a resolved entry was examined again: %+v", r2)
}
}
func TestRetry_OperatorToldOnceAfterThreeFailures(t *testing.T) {
e, _, _ := leakScratch(t)
var gave [][]int
for i := 0; i < MaxTeardownTries+2; i++ {
gave = append(gave, e.RetryScratchTeardown(context.Background()).GaveUp)
}
for i, g := range gave {
want := 0
if i == MaxTeardownTries-1 {
want = 1
}
if len(g) != want {
t.Fatalf("pass %d gave up on %v — want the operator told exactly once, on pass %d", i+1, g, MaxTeardownTries)
}
}
}
func TestRetry_NeverTouchesARunningTest(t *testing.T) {
e, api, _ := leakScratch(t)
api.destroyErr = nil
e.markScratch(990000, true) // a restore-test is (again) working on this vmid
before := len(api.destroys)
if r := e.RetryScratchTeardown(context.Background()); r.Examined != 0 || len(api.destroys) != before {
t.Fatalf("the timer touched a scratch a running test owns: %+v destroys=%v", r, api.destroys)
}
}
+176
View File
@@ -0,0 +1,176 @@
package reconcile
import (
"context"
"fmt"
"sort"
"strings"
)
// ── The restore-test's space preflight (R-672, agent v0.133.0) ─────────────────────────────────────
//
// MEASURED 2026-09-24 on demo-hp: the scheduled restore-test restored 9201's archive into `local-lvm`
// — the SAME thin pool that holds 9201 — with no free-space check. The pool reached 100 %
// (`out_of_data_space`, `error_if_no_space`), and 9201's rootfs and data volume remounted READ-ONLY.
// Evidence: felhom.eu `documentation/audits/night-2026-09-24/C-02…C-07`, `audits/r672-2026-09-24/`.
//
// THE THREE RULES, all decided here before ANY mutation (before the scratch entry is even journaled):
// 1. SPACE FIRST. The target storage must have free data ≥ restored × factor + reserve (defaults 1.2 and
// 5 GiB, `backup.restore_test_space_factor` / `backup.restore_test_space_reserve_gib`), and a thin
// pool's metadata must have room for the same share. `restored` is the UNCOMPRESSED size — the
// archive FILE is the wrong number: 9201's archive was 6.9 GB and its restore wrote 22.6 GB, so
// "file × 1.2 + 5 GiB" (14.3 GB) would have let the 2026-09-24 test run into a pool with 23 GB free.
// 2. KEEP OFF THE TESTED GUEST'S POOL when another eligible storage (active, takes `rootdir`, and the
// agent holds Datastore.AllocateSpace on it) passes rule 1. With only one, rule 1 decides.
// 3. UNKNOWN REFUSES. An unreadable size, an unreadable storage or an unknown thin-pool metadata fill is
// a skip with its reason, never a guess — the fail-safe direction of every guard in this project.
// A refusal is reported to the hub as the test's RESULT ("skipped: …", pass=false), never as a pass.
// Pinned by restoretest_space_test.go.
// RestoreSpace is the preflight's seam onto the host. Production: internal/restorespace.
type RestoreSpace interface {
// RestoredBytes is how many bytes restoring `archive` will write (uncompressed), and where that
// figure came from (for the log and the refusal).
RestoredBytes(ctx context.Context, archive string) (bytes int64, source string, err error)
// Free reports the storage's free data bytes and, for a thin pool, its metadata-used fraction.
Free(ctx context.Context, storage string) (StorageFree, error)
// Eligible lists the storages a restore-test may target: active, content `rootdir`, and the agent
// holds Datastore.AllocateSpace there.
Eligible(ctx context.Context) ([]string, error)
}
// StorageFree is one storage's free space as the preflight judges it.
type StorageFree struct {
AvailBytes int64
UsedBytes int64
Thin bool
// MetaUsedFraction is the thin pool's metadata use (0..1); MetaKnown false = could not be read.
MetaUsedFraction float64
MetaKnown bool
}
// SpacePolicy is rule 1's margin.
type SpacePolicy struct {
Factor float64 // ≥ 1
ReserveBytes int64
}
// DefaultSpacePolicy is 1.2 × restored + 5 GiB.
var DefaultSpacePolicy = SpacePolicy{Factor: 1.2, ReserveBytes: 5 << 30}
// SpaceVerdict is the preflight's answer.
type SpaceVerdict struct {
OK bool
Storage string // the storage the restore goes to (when OK) or was judged (when not)
Required int64
Avail int64
Reason string // empty when OK
// Avoided is the tested guest's own storage, when rule 2 moved the restore off it.
Avoided string
}
// requiredBytes is rule 1's figure.
func (p SpacePolicy) requiredBytes(restored int64) int64 {
f := p.Factor
if f < 1 {
f = DefaultSpacePolicy.Factor
}
return int64(float64(restored)*f) + p.ReserveBytes
}
// fits judges one storage against rule 1 (data AND thin metadata). An unknown metadata fill on a thin
// pool refuses (rule 3).
func fits(fr StorageFree, required int64) (bool, string) {
if fr.AvailBytes < required {
return false, fmt.Sprintf("needs %s free, has %s", gib(required), gib(fr.AvailBytes))
}
if fr.Thin {
if !fr.MetaKnown {
return false, "thin-pool metadata fill unknown"
}
// The metadata a restore of `required` bytes needs, in the pool's own proportion of metadata to
// data. A pool with no data yet has no proportion to read → only the absolute ceiling applies.
need := 0.0
if fr.UsedBytes > 0 {
need = fr.MetaUsedFraction * float64(required) / float64(fr.UsedBytes)
}
if fr.MetaUsedFraction+need > 0.9 {
return false, fmt.Sprintf("thin-pool metadata would reach %.0f%% (now %.0f%%)", 100*(fr.MetaUsedFraction+need), 100*fr.MetaUsedFraction)
}
}
return true, ""
}
func gib(b int64) string { return fmt.Sprintf("%.1f GiB", float64(b)/(1<<30)) }
// sourceStorages returns the storage ids that hold the ARCHIVED guest's volumes (rootfs and every mpN
// that names a `storage:volume`), read from the archive's own embedded config — the guest under test.
// Bind mounts (a leading "/") carry no storage.
func sourceStorages(rawCfg string) map[string]bool {
out := map[string]bool{}
for k, v := range archiveCurrentConfig(rawCfg) {
if k != "rootfs" && !(strings.HasPrefix(k, "mp") && len(k) > 2 && k[2] >= '0' && k[2] <= '9') {
continue
}
vol := strings.TrimSpace(strings.SplitN(strings.TrimSpace(v), ",", 2)[0])
if vol == "" || strings.HasPrefix(vol, "/") {
continue
}
if st, _, ok := strings.Cut(vol, ":"); ok && st != "" {
out[st] = true
}
}
return out
}
// PreflightRestoreSpace applies the three rules. `configured` is `backup.restore_storage`.
func PreflightRestoreSpace(ctx context.Context, space RestoreSpace, policy SpacePolicy, archive, rawCfg, configured string) SpaceVerdict {
if space == nil {
return SpaceVerdict{Storage: configured, Reason: "no space check is wired — refusing (fail-closed)"}
}
restored, src, err := space.RestoredBytes(ctx, archive)
if err != nil || restored <= 0 {
return SpaceVerdict{Storage: configured, Reason: fmt.Sprintf("cannot tell how much the restore writes (%v)", err)}
}
required := policy.requiredBytes(restored)
own := sourceStorages(rawCfg)
// Rule 2: the configured storage holds the guest under test → try the others first.
var order []string
avoided := ""
if own[configured] {
eligible, eerr := space.Eligible(ctx)
if eerr == nil {
sort.Strings(eligible)
for _, s := range eligible {
if s != configured && !own[s] {
order = append(order, s)
}
}
}
if len(order) > 0 {
avoided = configured
}
}
order = append(order, configured)
var last SpaceVerdict
for _, s := range order {
fr, ferr := space.Free(ctx, s)
if ferr != nil {
last = SpaceVerdict{Storage: s, Required: required, Reason: fmt.Sprintf("cannot read free space on %s (%v)", s, ferr)}
continue
}
ok, why := fits(fr, required)
v := SpaceVerdict{OK: ok, Storage: s, Required: required, Avail: fr.AvailBytes}
if ok {
if s != configured {
v.Avoided = avoided
}
return v
}
v.Reason = fmt.Sprintf("not enough space on %s: restoring %s (%s) %s", s, gib(restored), src, why)
last = v
}
return last
}
@@ -0,0 +1,156 @@
package reconcile
import (
"context"
"errors"
"path/filepath"
"strings"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-672 (v0.133.0) — the restore-test's space preflight. Every test asserts the CONSEQUENCE: whether the
// Proxmox API was asked to restore anything, where to, and what the result says — never only the verdict.
const gb = int64(1000 * 1000 * 1000)
// fakeSpace is a configurable RestoreSpace.
type fakeSpace struct {
restored int64
restoredErr error
free map[string]StorageFree
freeErr map[string]error
eligible []string
}
func (f fakeSpace) RestoredBytes(context.Context, string) (int64, string, error) {
return f.restored, "fake", f.restoredErr
}
func (f fakeSpace) Free(_ context.Context, s string) (StorageFree, error) {
if err := f.freeErr[s]; err != nil {
return StorageFree{}, err
}
fr, ok := f.free[s]
if !ok {
return StorageFree{}, errors.New("unknown storage")
}
return fr, nil
}
func (f fakeSpace) Eligible(context.Context) ([]string, error) { return f.eligible, nil }
// thin9201 is demo-hp's local-lvm at 10:29 on 2026-09-24, just before the restore-test that filled it:
// 23.2 GB free, 33.3 GB used, metadata 2.65 %.
var thin9201 = StorageFree{AvailBytes: 23210892 * 1024, UsedBytes: 33277043 * 1024, Thin: true, MetaUsedFraction: 0.0265, MetaKnown: true}
// archive9201 is 9201's archive config: both volumes on local-lvm.
const archive9201 = "hostname: demo-hp\nrootfs: local-lvm:vm-9201-disk-0,size=32G\nmp0: local-lvm:vm-9201-disk-1,mp=/var/lib/felhom,backup=1,size=70G\nmp8: /mnt/felhom-drives,mp=/mnt/felhom-drives\n"
func spaceEngine(t *testing.T, api *fakeAPI, sp RestoreSpace) (*Engine, *Journal) {
t.Helper()
j, err := OpenJournal(filepath.Join(t.TempDir(), "journal.log"))
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { j.Close() })
q := NewQueue()
t.Cleanup(q.Close)
return NewEngine(EngineOptions{API: api, Queue: q, Journal: j, RestoreSpace: sp}), j
}
func run9201(e *Engine) RestoreTestResult {
return e.RunRestoreTest(context.Background(), RestoreTestSpec{
Archive: "local:backup/vzdump-lxc-9201-2026_09_23-06_55_25.tar.zst", RestoreStorage: "local-lvm",
ScratchMin: 990000, ScratchMax: 990009, SourceTier: "local",
})
}
// TestSpace_The2026_09_24TestIsRefused replays R-672: 9201's archive restores 22.6 GB (its vzdump log),
// the pool has 23.2 GB free. The test must NOT start — no restore call, no journaled scratch — and the
// result must say why, as a non-pass.
//
// COMPANION RED-PROOFS (REPORT): (1) the preflight removed (v0.132.0's shape) → a restore into local-lvm
// is issued; (2) `restored` taken from the archive FILE (6.9 GB, the brief's "archive × 1.2 + 5 GiB") →
// 6.9×1.2+5.4 = 13.7 GB < 23.2 GB free, so the test starts — the defect the uncompressed size exists for.
func TestSpace_The2026_09_24TestIsRefused(t *testing.T) {
api := &fakeAPI{extractCfg: archive9201, cfg: map[int]proxmox.GuestConfig{990000: scratchCfg()}}
e, j := spaceEngine(t, api, fakeSpace{restored: 22607360000, free: map[string]StorageFree{"local-lvm": thin9201}})
res := run9201(e)
if len(api.restores) != 0 {
t.Fatalf("a restore was issued into a pool that cannot take it: %+v", api.restores)
}
if len(j.InFlight()) != 0 {
t.Fatalf("a scratch entry was journaled for a test that must not start: %+v", j.InFlight())
}
if res.Pass || !res.Skipped || !strings.Contains(res.SkipReason, "not enough space on local-lvm") {
t.Fatalf("result = pass=%v skipped=%v reason=%q — want a non-pass skip naming the storage", res.Pass, res.Skipped, res.SkipReason)
}
if res.RequiredBytes < 32*gb || res.AvailBytes != thin9201.AvailBytes {
t.Fatalf("required=%d avail=%d — want ≥ 32 GB required (22.6 × 1.2 + 5 GiB) against 23.2 GB", res.RequiredBytes, res.AvailBytes)
}
}
// TestSpace_KeepsOffTheTestedGuestsPool — rule 2: another eligible storage that fits takes the restore.
func TestSpace_KeepsOffTheTestedGuestsPool(t *testing.T) {
api := &fakeAPI{extractCfg: archive9201, cfg: map[int]proxmox.GuestConfig{990000: scratchCfg()}}
e, _ := spaceEngine(t, api, fakeSpace{restored: 22607360000, eligible: []string{"local-lvm", "big-dir"},
free: map[string]StorageFree{"local-lvm": {AvailBytes: 900 * gb, UsedBytes: 10 * gb, Thin: true, MetaKnown: true}, "big-dir": {AvailBytes: 500 * gb}}})
res := run9201(e)
if len(api.restores) != 1 || api.restores[0].Storage != "big-dir" {
t.Fatalf("restores = %+v — want ONE restore onto big-dir, off 9201's own pool", api.restores)
}
if res.TargetStorage != "big-dir" {
t.Fatalf("target=%q", res.TargetStorage)
}
for k, v := range api.restores[0].MountOverrides {
if strings.HasPrefix(v, "local-lvm:") {
t.Fatalf("%s still lands on the tested guest's pool: %s", k, v)
}
}
}
// TestSpace_OnlyOnePool_RuleOneDecides — no other eligible storage: the tested guest's pool is used when it
// fits (demo-hp's real shape: nvme-scratch takes rootdir but the agent holds no AllocateSpace there).
func TestSpace_OnlyOnePool_RuleOneDecides(t *testing.T) {
api := &fakeAPI{extractCfg: archive9201, cfg: map[int]proxmox.GuestConfig{990000: scratchCfg()}}
e, _ := spaceEngine(t, api, fakeSpace{restored: 2 * gb, eligible: []string{"local-lvm"}, free: map[string]StorageFree{"local-lvm": thin9201}})
res := run9201(e)
if len(api.restores) != 1 || api.restores[0].Storage != "local-lvm" || res.Skipped || res.TargetStorage != "local-lvm" {
t.Fatalf("restores=%+v skipped=%v — a 2 GB restore fits 23 GB free on the only pool", api.restores, res.Skipped)
}
}
// TestSpace_UnknownRefuses — rule 3, one case per unknown. Nothing is restored in any of them.
func TestSpace_UnknownRefuses(t *testing.T) {
cases := map[string]RestoreSpace{
"no space check wired": nil,
"restore size unknown": fakeSpace{restoredErr: errors.New("no vzdump log"), free: map[string]StorageFree{"local-lvm": thin9201}},
"free space unreadable": fakeSpace{restored: gb, freeErr: map[string]error{"local-lvm": errors.New("api down")}},
"thin metadata unknown": fakeSpace{restored: gb, free: map[string]StorageFree{"local-lvm": {AvailBytes: 900 * gb, UsedBytes: gb, Thin: true}}},
"metadata would overrun": fakeSpace{restored: 10 * gb, free: map[string]StorageFree{"local-lvm": {AvailBytes: 900 * gb, UsedBytes: 10 * gb, Thin: true, MetaUsedFraction: 0.5, MetaKnown: true}}},
}
for name, sp := range cases {
t.Run(name, func(t *testing.T) {
api := &fakeAPI{extractCfg: archive9201, cfg: map[int]proxmox.GuestConfig{990000: scratchCfg()}}
var e *Engine
if sp == nil {
e, _ = spaceEngine(t, api, nil)
} else {
e, _ = spaceEngine(t, api, sp)
}
res := run9201(e)
if len(api.restores) != 0 || res.Pass || !res.Skipped || res.SkipReason == "" {
t.Fatalf("restores=%d pass=%v skipped=%v reason=%q — an unknown must refuse before anything moves",
len(api.restores), res.Pass, res.Skipped, res.SkipReason)
}
})
}
}
// TestSpace_SourceStorages reads the tested guest's pools from the ARCHIVE's config, binds excluded.
func TestSpace_SourceStorages(t *testing.T) {
got := sourceStorages(archive9201 + "mp1: other:vm-9201-disk-2,mp=/x,size=1G\n[snap]\nrootfs: snapstore:x\n")
if !got["local-lvm"] || !got["other"] || got["snapstore"] || len(got) != 2 {
t.Fatalf("sourceStorages = %v — want local-lvm + other, binds and snapshot sections excluded", got)
}
}
+16
View File
@@ -0,0 +1,16 @@
package reconcile
import "context"
// roomySpace is the permissive RestoreSpace the pre-R-672 restore-test tests run with: 1 GiB restored,
// 1 TiB free, thin metadata known and low. The space rules themselves are pinned in
// restoretest_space_test.go.
type roomySpace struct{}
func (roomySpace) RestoredBytes(context.Context, string) (int64, string, error) {
return 1 << 30, "test", nil
}
func (roomySpace) Free(context.Context, string) (StorageFree, error) {
return StorageFree{AvailBytes: 1 << 40, UsedBytes: 1 << 30, Thin: true, MetaUsedFraction: 0.01, MetaKnown: true}, nil
}
func (roomySpace) Eligible(context.Context) ([]string, error) { return nil, nil }
+176
View File
@@ -0,0 +1,176 @@
// Package restorespace is the production seam behind reconcile.RestoreSpace (R-672, agent v0.133.0):
// how much a restore of an archive writes, how much a storage has free, and which storages a
// restore-test may target. Every read that cannot answer returns an error — the preflight then
// REFUSES (reconcile/restoretest_space.go rule 3); nothing here guesses.
package restorespace
import (
"context"
"fmt"
"os"
"path"
"regexp"
"strconv"
"strings"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
"gitea.dooplex.hu/admin/felhom-agent/internal/reconcile"
)
// API is the Proxmox subset the provider reads.
type API interface {
ListStorage(ctx context.Context) ([]proxmox.Storage, error)
NodeStorage(ctx context.Context) ([]proxmox.Storage, error)
StorageContent(ctx context.Context, store string) ([]proxmox.StorageContent, error)
Permissions(ctx context.Context, aclPath string) (map[string]int, error)
}
// Provider implements reconcile.RestoreSpace.
type Provider struct {
API API
// ThinMeta reads a thin pool's metadata-used fraction (storage.HostOps.ThinPoolMetadata).
ThinMeta func(ctx context.Context, vg, pool string) (float64, bool)
// ReadFile reads a vzdump log; nil → os.ReadFile.
ReadFile func(name string) ([]byte, error)
}
var _ reconcile.RestoreSpace = (*Provider)(nil)
// totalWrittenRe is vzdump's own count of the bytes tar wrote into the archive — the UNCOMPRESSED size,
// i.e. what a restore writes back ("INFO: Total bytes written: 22607360000 (22GiB, 49MiB/s)").
var totalWrittenRe = regexp.MustCompile(`Total bytes written:\s*(\d+)`)
// archiveExts are the vzdump archive suffixes; the log is the archive name without it + ".log".
var archiveExts = []string{".tar.zst", ".tar.gz", ".tar.lzo", ".tgz", ".tar"}
func (p *Provider) readFile(name string) ([]byte, error) {
if p.ReadFile != nil {
return p.ReadFile(name)
}
return os.ReadFile(name)
}
// RestoredBytes: a file-backed archive → its vzdump log's "Total bytes written"; a PBS archive → the
// size Proxmox reports for the snapshot (its logical, uncompressed size). The archive FILE size is never
// used: it is compressed (6.9 GB for a 22.6 GB restore, measured 2026-09-24).
func (p *Provider) RestoredBytes(ctx context.Context, archive string) (int64, string, error) {
id, vol, ok := strings.Cut(archive, ":")
if !ok || id == "" || vol == "" {
return 0, "", fmt.Errorf("not a storage volid: %q", archive)
}
st, err := p.storageConfig(ctx, id)
if err != nil {
return 0, "", err
}
switch st.Type {
case "pbs":
items, err := p.API.StorageContent(ctx, id)
if err != nil {
return 0, "", fmt.Errorf("list %s: %w", id, err)
}
for _, it := range items {
if it.VolID == archive && it.Size > 0 {
return it.Size, "pbs snapshot size", nil
}
}
return 0, "", fmt.Errorf("archive %s not listed on %s with a size", archive, id)
default:
if st.Path == "" {
return 0, "", fmt.Errorf("storage %s (%s) has no path to read a vzdump log from", id, st.Type)
}
base := path.Base(vol) // "backup/vzdump-lxc-…tar.zst" → "vzdump-lxc-…tar.zst"
stem := ""
for _, ext := range archiveExts {
if strings.HasSuffix(base, ext) {
stem = strings.TrimSuffix(base, ext)
break
}
}
if stem == "" {
return 0, "", fmt.Errorf("unknown archive suffix: %s", base)
}
logPath := path.Join(st.Path, "dump", stem+".log")
b, err := p.readFile(logPath)
if err != nil {
return 0, "", fmt.Errorf("read vzdump log %s: %w", logPath, err)
}
m := totalWrittenRe.FindSubmatch(b)
if m == nil {
return 0, "", fmt.Errorf("vzdump log %s carries no \"Total bytes written\"", logPath)
}
n, err := strconv.ParseInt(string(m[1]), 10, 64)
if err != nil || n <= 0 {
return 0, "", fmt.Errorf("vzdump log %s: bad byte count %q", logPath, m[1])
}
return n, "vzdump log: total bytes written", nil
}
}
func (p *Provider) storageConfig(ctx context.Context, id string) (proxmox.Storage, error) {
all, err := p.API.ListStorage(ctx)
if err != nil {
return proxmox.Storage{}, fmt.Errorf("list storage config: %w", err)
}
for _, s := range all {
if s.Storage == id {
return s, nil
}
}
return proxmox.Storage{}, fmt.Errorf("storage %s not configured", id)
}
// Free reads the node's live usage for `storage`; a thin pool adds its metadata fill.
func (p *Provider) Free(ctx context.Context, storage string) (reconcile.StorageFree, error) {
live, err := p.API.NodeStorage(ctx)
if err != nil {
return reconcile.StorageFree{}, fmt.Errorf("node storage: %w", err)
}
for _, s := range live {
if s.Storage != storage {
continue
}
if s.Active != 1 {
return reconcile.StorageFree{}, fmt.Errorf("storage %s is not active", storage)
}
fr := reconcile.StorageFree{AvailBytes: s.Avail, UsedBytes: s.Used, Thin: s.Type == "lvmthin"}
if fr.Thin {
cfg, cerr := p.storageConfig(ctx, storage)
if cerr == nil && p.ThinMeta != nil && cfg.VGName != "" && cfg.ThinPool != "" {
fr.MetaUsedFraction, fr.MetaKnown = p.ThinMeta(ctx, cfg.VGName, cfg.ThinPool)
}
}
return fr, nil
}
return reconcile.StorageFree{}, fmt.Errorf("storage %s not reported by the node", storage)
}
// Eligible: active, content includes `rootdir`, and the agent holds Datastore.AllocateSpace on
// /storage/<id>. The permission is read for the SPECIFIC privilege — the box-wide grant answers every
// path with inherited privileges (proxmox.Client.Permissions).
func (p *Provider) Eligible(ctx context.Context) ([]string, error) {
live, err := p.API.NodeStorage(ctx)
if err != nil {
return nil, fmt.Errorf("node storage: %w", err)
}
var out []string
for _, s := range live {
if s.Active != 1 || !hasContent(s.Content, "rootdir") {
continue
}
privs, perr := p.API.Permissions(ctx, "/storage/"+s.Storage)
if perr != nil || privs["Datastore.AllocateSpace"] != 1 {
continue
}
out = append(out, s.Storage)
}
return out, nil
}
func hasContent(list, want string) bool {
for _, c := range strings.Split(list, ",") {
if strings.TrimSpace(c) == want {
return true
}
}
return false
}
+133
View File
@@ -0,0 +1,133 @@
package restorespace
import (
"context"
"errors"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-672 (v0.133.0). The provider behind the restore-test's space preflight. No test reaches a real
// Proxmox or a real file: the API and ReadFile are fakes.
type fakeAPI struct {
cfg []proxmox.Storage
live []proxmox.Storage
content map[string][]proxmox.StorageContent
perms map[string]map[string]int
}
func (f fakeAPI) ListStorage(context.Context) ([]proxmox.Storage, error) { return f.cfg, nil }
func (f fakeAPI) NodeStorage(context.Context) ([]proxmox.Storage, error) { return f.live, nil }
func (f fakeAPI) StorageContent(_ context.Context, s string) ([]proxmox.StorageContent, error) {
return f.content[s], nil
}
func (f fakeAPI) Permissions(_ context.Context, p string) (map[string]int, error) {
if m, ok := f.perms[p]; ok {
return m, nil
}
return map[string]int{}, nil
}
const archive = "local:backup/vzdump-lxc-9201-2026_09_23-06_55_25.tar.zst"
// The real log's tail (demo-hp, 2026-09-23): 22.6 GB written, a 6.91 GB archive file.
const vzdumpLog = "2026-09-23 07:02:47 INFO: Total bytes written: 22607360000 (22GiB, 49MiB/s)\n2026-09-23 07:02:47 INFO: archive file size: 6.91GB\n"
// demoHP is demo-hp's storage layout: `local` (dir, backups), `local-lvm` (thin), `nvme-scratch` (dir,
// rootdir — but the agent holds NO grant there), a pbs.
func demoHP() fakeAPI {
return fakeAPI{
cfg: []proxmox.Storage{
{Storage: "local", Type: "dir", Path: "/var/lib/vz"},
{Storage: "local-lvm", Type: "lvmthin", VGName: "pve", ThinPool: "data"},
{Storage: "nvme-scratch", Type: "dir", Path: "/mnt/hdd_1"},
{Storage: "felhom-pbs", Type: "pbs"},
},
live: []proxmox.Storage{
{Storage: "local", Type: "dir", Content: "vztmpl,backup,iso,import", Active: 1, Avail: 4 << 30, Used: 34 << 30},
{Storage: "local-lvm", Type: "lvmthin", Content: "images,rootdir", Active: 1, Avail: 23210892 * 1024, Used: 33277043 * 1024},
{Storage: "nvme-scratch", Type: "dir", Content: "images,rootdir", Active: 1, Avail: 800 << 30},
{Storage: "felhom-pbs", Type: "pbs", Content: "backup", Active: 0},
},
content: map[string][]proxmox.StorageContent{
"local": {{VolID: archive, Size: 7417540996}},
"felhom-pbs": {{VolID: "felhom-pbs:backup/ct/9201/2026-09-23T02:00:00Z", Size: 21 << 30}},
},
perms: map[string]map[string]int{
"/storage/local": {"Datastore.Audit": 1, "Datastore.AllocateSpace": 1},
"/storage/local-lvm": {"Datastore.Audit": 1, "Datastore.AllocateSpace": 1},
// nvme-scratch: only the inherited box-wide Datastore.Audit — the trap proxmox.Permissions names.
"/storage/nvme-scratch": {"Datastore.Audit": 1},
},
}
}
// TestRestoredBytes_ReadsTheUncompressedSize — the vzdump log's "Total bytes written", never the archive
// FILE size (6.9 GB for a 22.6 GB restore).
//
// COMPANION RED-PROOF (REPORT): return the storage content's Size for a dir storage → 7417540996, and
// this test fails at "the compressed file size was used".
func TestRestoredBytes_ReadsTheUncompressedSize(t *testing.T) {
var asked string
p := &Provider{API: demoHP(), ReadFile: func(n string) ([]byte, error) { asked = n; return []byte(vzdumpLog), nil }}
n, src, err := p.RestoredBytes(context.Background(), archive)
if err != nil {
t.Fatal(err)
}
if n == 7417540996 {
t.Fatal("the compressed file size was used — the restore writes 3× that")
}
if n != 22607360000 || asked != "/var/lib/vz/dump/vzdump-lxc-9201-2026_09_23-06_55_25.log" {
t.Fatalf("n=%d from %q (log %q)", n, src, asked)
}
}
func TestRestoredBytes_UnknownIsAnError(t *testing.T) {
cases := map[string]*Provider{
"no log": {API: demoHP(), ReadFile: func(string) ([]byte, error) { return nil, errors.New("ENOENT") }},
"log without size": {API: demoHP(), ReadFile: func(string) ([]byte, error) { return []byte("ERROR: failed\n"), nil }},
}
for name, p := range cases {
if n, _, err := p.RestoredBytes(context.Background(), archive); err == nil {
t.Fatalf("%s: got %d, want an error (the preflight then refuses)", name, n)
}
}
if _, _, err := (&Provider{API: demoHP()}).RestoredBytes(context.Background(), "local:backup/weird.vma"); err == nil {
t.Fatal("an unknown archive suffix must be an error")
}
}
func TestRestoredBytes_PBS(t *testing.T) {
p := &Provider{API: demoHP()}
n, _, err := p.RestoredBytes(context.Background(), "felhom-pbs:backup/ct/9201/2026-09-23T02:00:00Z")
if err != nil || n != 21<<30 {
t.Fatalf("n=%d err=%v", n, err)
}
}
// TestEligible_NeedsTheSpecificGrant — nvme-scratch takes rootdir but the agent holds only the inherited
// Datastore.Audit there, so it is NOT eligible; `local` holds no rootdir.
func TestEligible_NeedsTheSpecificGrant(t *testing.T) {
got, err := (&Provider{API: demoHP()}).Eligible(context.Background())
if err != nil || len(got) != 1 || got[0] != "local-lvm" {
t.Fatalf("eligible = %v (%v) — want only local-lvm on demo-hp", got, err)
}
}
func TestFree_ThinCarriesMetadata(t *testing.T) {
p := &Provider{API: demoHP(), ThinMeta: func(_ context.Context, vg, pool string) (float64, bool) {
if vg != "pve" || pool != "data" {
t.Fatalf("metadata read for %s/%s", vg, pool)
}
return 0.0265, true
}}
fr, err := p.Free(context.Background(), "local-lvm")
if err != nil || !fr.Thin || !fr.MetaKnown || fr.MetaUsedFraction != 0.0265 || fr.AvailBytes != 23210892*1024 {
t.Fatalf("free = %+v err=%v", fr, err)
}
if _, err := p.Free(context.Background(), "felhom-pbs"); err == nil {
t.Fatal("an inactive storage must be an error")
}
}
+89 -2
View File
@@ -53,7 +53,11 @@ type claimFacts struct {
nodes []claimNode // the whole disk + its children (partitions)
lvmPV bool // pvs (authoritative): the disk / a partition is an LVM physical volume
zfsMember bool // zpool (authoritative): the disk / a partition is a ZFS pool member
gatherErr string // non-empty ⇒ a REQUIRED read failed ⇒ fail-safe CLAIMED
// felhomOwnedMounts (R-220) — mountpoints OUTSIDE /mnt/felhom-drives that are nevertheless Felhom's
// OWN, corroborated from the host mount table: the same device is also mounted at the managed path.
// Empty means "nothing corroborated", which is the fail-safe direction.
felhomOwnedMounts map[string]bool
gatherErr string // non-empty ⇒ a REQUIRED read failed ⇒ fail-safe CLAIMED
}
// classifyClaim is the pure guard verdict. unclaimed=true ONLY when the device is provably free for
@@ -81,7 +85,20 @@ func classifyClaim(f claimFacts) (unclaimed bool, reason string) {
if memberFSTypes[n.fstype] {
return false, "device holds a " + n.fstype + " (" + n.name + ")"
}
if n.mountpoint != "" && !underFelhomDrives(n.mountpoint) {
// ── R-220 — A MOUNT FELHOM ITSELF MADE IS NOT "SOMETHING ELSE". ───────────────────────
//
// Enrolment mounts a drive TWICE: at the managed path `/mnt/felhom-drives/<name>` and at the
// raw `/mnt/<name>` it creates on the host. The host — and therefore that raw mount — survives
// a guest rebuild, while the controller's registry does not. So after a rebuild the customer's
// own drives looked foreign, `attach` returned an empty list, and the refusal told them to
// choose from it. Measured live three times (CAMPAIGN-11 Phase 1, and the R-201 re-walk twice);
// unmounting only the raw mounts flipped `attach: []` to both drives every time.
//
// The fence this must NOT breach: a disk genuinely in use by something else stays refused. So
// the exemption is not "any /mnt/* path" — it is CORROBORATED: the same device must ALSO be
// mounted at Felhom's managed path, which is a state only Felhom's own enrolment produces.
// A foreign disk at /srv/data or /media/x has no such counterpart and is still refused.
if n.mountpoint != "" && !underFelhomDrives(n.mountpoint) && !f.felhomOwnedMounts[n.mountpoint] {
return false, "device is mounted at " + n.mountpoint + " (" + n.name + ")"
}
}
@@ -154,6 +171,8 @@ func (h *SudoHostOps) gatherClaimFacts(ctx context.Context, device string) claim
return f
}
f.nodes = nodes
// R-220: corroborate which non-managed mountpoints are nevertheless Felhom's own.
f.felhomOwnedMounts = felhomOwnedMounts(device, nodes, h.mountTable)
// LVM PV (authoritative). pvs installed but erroring ⇒ fail-safe claimed; absent ⇒ rely on lsblk's
// LVM2_member FSTYPE (already in nodes).
@@ -285,3 +304,71 @@ func (h *SudoHostOps) zfsMembers(ctx context.Context, nodes []claimNode, wholeDi
}
return false, nil
}
// mountTableSource yields the host mount table as (device, mountpoint) pairs. A seam so the R-220
// corroboration is unit-testable without a host. nil ⇒ the real /proc/mounts.
type mountTableSource func() ([][2]string, error)
// procMounts reads /proc/mounts — WORLD-READABLE, so this needs no sudo and no allowlisted command.
// That matters: the lsblk invocation is pinned verbatim in the sudoers file
// (`lsblk -J -o NAME,FSTYPE,PTTYPE,MOUNTPOINT /dev/*`), so switching it to the plural MOUNTPOINTS
// would have meant shipping a sudoers change with the binary — a far larger blast radius than this
// finding warrants. Reading the mount table directly sidesteps that entirely.
func procMounts() ([][2]string, error) {
data, err := os.ReadFile("/proc/mounts")
if err != nil {
return nil, err
}
var out [][2]string
for _, line := range strings.Split(string(data), "\n") {
fields := strings.Fields(line)
if len(fields) < 2 {
continue
}
// /proc/mounts escapes spaces as \040; unescape so a path with a space still compares.
out = append(out, [2]string{fields[0], strings.ReplaceAll(fields[1], `\040`, " ")})
}
return out, nil
}
// felhomOwnedMounts returns the mountpoints of `device` (and its children) that sit OUTSIDE
// /mnt/felhom-drives but are still Felhom's own, corroborated by the same device also being mounted
// UNDER /mnt/felhom-drives. That pairing is what enrolment produces and nothing else does.
//
// ⚠ FAIL-SAFE: an unreadable mount table returns an EMPTY set, never a permissive one. The device then
// classifies exactly as it did before R-220 — refused — because "we could not corroborate" must never
// read as "it is ours".
func felhomOwnedMounts(device string, nodes []claimNode, src mountTableSource) map[string]bool {
if src == nil {
src = procMounts
}
table, err := src()
if err != nil {
return nil // unreadable ⇒ corroborate nothing
}
// Every device name this disk answers to: the whole disk and each child node.
devs := map[string]bool{device: true}
if wd, ok := wholeDiskOf(device); ok {
devs[wd] = true
}
for _, n := range nodes {
devs["/dev/"+n.name] = true
}
// A device is Felhom-managed only if it is mounted under the managed prefix.
managed := map[string]bool{}
for _, row := range table {
if devs[row[0]] && underFelhomDrives(row[1]) {
managed[row[0]] = true
}
}
if len(managed) == 0 {
return nil
}
owned := map[string]bool{}
for _, row := range table {
if managed[row[0]] && !underFelhomDrives(row[1]) {
owned[path.Clean(row[1])] = true
}
}
return owned
}
+102
View File
@@ -0,0 +1,102 @@
package storage
import "testing"
// ── R-220 — A MOUNT FELHOM ITSELF MADE IS NOT "SOMETHING ELSE" ──────────────────────────────────
//
// Enrolment mounts a drive twice: at `/mnt/felhom-drives/<name>` and at the raw `/mnt/<name>` it
// creates on the host. The host survives a guest rebuild; the controller's registry does not. So after
// a rebuild the customer's own drives read as claimed-by-something-else, `attach` came back empty, and
// the refusal told them to pick from the empty list. Measured three times live.
//
// The fence: a disk genuinely in use elsewhere must STILL be refused. These assert both directions.
// ── SCENARIO E — the customer's own drive is offered again after a rebuild ───────────────────────
//
// RED-PROOF: drop `&& !f.felhomOwnedMounts[n.mountpoint]` from classifyClaim — the pre-R-220 check —
// and this FAILS with the drive refused and the list empty again.
func TestClassifyClaim_R220_FelhomsOwnRawMountIsNotForeign(t *testing.T) {
f := claimFacts{
device: "/dev/sdb", wholeDisk: "/dev/sdb", wholeDiskOK: true,
nodes: []claimNode{{name: "sdb", fstype: "ext4", mountpoint: "/mnt/adatok"}},
// corroborated: the SAME device is also mounted at the managed path
felhomOwnedMounts: map[string]bool{"/mnt/adatok": true},
}
unclaimed, reason := classifyClaim(f)
if !unclaimed {
t.Fatalf("R-220 RETURNED: the customer's own drive is refused after a rebuild — %q", reason)
}
}
// ── SCENARIO F — a genuinely foreign mount is STILL refused ──────────────────────────────────────
//
// RED-PROOF: over-widen the fix to exempt any /mnt/* path (or to skip the mountpoint check entirely)
// and this FAILS — a disk another system is using would be offered for formatting.
func TestClassifyClaim_R220_ForeignMountIsStillRefused(t *testing.T) {
for _, mp := range []string{"/srv/data", "/media/photos", "/mnt/someone-elses-disk", "/var/lib/other"} {
f := claimFacts{
device: "/dev/sdb", wholeDisk: "/dev/sdb", wholeDiskOK: true,
nodes: []claimNode{{name: "sdb", fstype: "ext4", mountpoint: mp}},
felhomOwnedMounts: nil, // nothing corroborated it as ours
}
unclaimed, reason := classifyClaim(f)
if unclaimed {
t.Fatalf("THE FENCE BROKE: a disk mounted at %s was offered for formatting", mp)
}
if reason == "" {
t.Fatalf("a refusal must carry a reason (%s)", mp)
}
}
}
// The corroboration itself: it must require BOTH mounts of the SAME device, and fail safe.
func TestFelhomOwnedMounts_RequiresTheManagedCounterpart(t *testing.T) {
nodes := []claimNode{{name: "sdb"}}
t.Run("both mounts present -> the raw one is ours", func(t *testing.T) {
src := func() ([][2]string, error) {
return [][2]string{
{"/dev/sdb", "/mnt/adatok"},
{"/dev/sdb", "/mnt/felhom-drives/adatok"},
}, nil
}
got := felhomOwnedMounts("/dev/sdb", nodes, src)
if !got["/mnt/adatok"] {
t.Fatal("the raw enrolment mount was not recognised as Felhom's own")
}
})
t.Run("only the raw mount -> corroborates NOTHING", func(t *testing.T) {
src := func() ([][2]string, error) {
return [][2]string{{"/dev/sdb", "/mnt/adatok"}}, nil
}
if got := felhomOwnedMounts("/dev/sdb", nodes, src); len(got) != 0 {
t.Fatalf("a lone /mnt/<name> mount must corroborate nothing, got %v", got)
}
})
t.Run("a DIFFERENT device under the managed path does not vouch for this one", func(t *testing.T) {
src := func() ([][2]string, error) {
return [][2]string{
{"/dev/sdb", "/srv/data"},
{"/dev/sdc", "/mnt/felhom-drives/mentes"}, // someone else's, not sdb's
}, nil
}
if got := felhomOwnedMounts("/dev/sdb", nodes, src); got["/srv/data"] {
t.Fatal("another device's managed mount vouched for a foreign one")
}
})
t.Run("an unreadable mount table corroborates NOTHING (fail-safe)", func(t *testing.T) {
src := func() ([][2]string, error) { return nil, errRead }
if got := felhomOwnedMounts("/dev/sdb", nodes, src); len(got) != 0 {
t.Fatalf("an unreadable mount table must corroborate nothing, got %v", got)
}
})
}
var errRead = errNoMountTable{}
type errNoMountTable struct{}
func (errNoMountTable) Error() string { return "mount table unreadable" }
+3
View File
@@ -164,6 +164,9 @@ type SudoHostOps struct {
// UNPRIVILEGED read (`systemctl is-failed`) — seam-injected so the reassert's F10 reset-failed path
// is unit-testable without a real systemd. Default set in NewSudoHostOps.
unitFailed func(ctx context.Context, unit string) bool
// mountTable (R-220) yields the host mount table for the "is this mount Felhom's own?"
// corroboration. nil ⇒ the real /proc/mounts; tests inject.
mountTable mountTableSource
}
// SudoHostOpsConfig configures a SudoHostOps.
+41
View File
@@ -6,6 +6,7 @@ import (
"log/slog"
"regexp"
"strings"
"sync"
"gitea.dooplex.hu/admin/felhom-agent/internal/hub"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
@@ -35,6 +36,45 @@ type Observer struct {
host HostReader
ops HostOps
logger *slog.Logger
// R-672 (v0.133.0): a thin pool crossing thinPoolAlarmFraction (data OR metadata) requests an
// out-of-band host report at once, so the hub's storage-fill alarm sees it in seconds instead of at
// the next 15-minute report. Rising edge per pool; re-armed below thinPoolRearmFraction.
highMu sync.Mutex
high map[string]bool
onThinHigh func()
}
const (
thinPoolAlarmFraction = 0.90
thinPoolRearmFraction = 0.85
)
// SetThinHighTrigger wires the out-of-band report request (main: the storage trigger channel).
func (o *Observer) SetThinHighTrigger(f func()) { o.onThinHigh = f }
// noteThinFill is the edge detector. key separates data from metadata so each has its own edge.
func (o *Observer) noteThinFill(key string, frac float64) {
o.highMu.Lock()
if o.high == nil {
o.high = map[string]bool{}
}
fire := false
switch {
case frac >= thinPoolAlarmFraction && !o.high[key]:
o.high[key] = true
fire = true
case frac < thinPoolRearmFraction && o.high[key]:
o.high[key] = false
}
o.highMu.Unlock()
if fire {
o.logger.Error("storage: thin pool crossed 90% — requesting an immediate host report (the hub alarms)",
"pool", key, "fraction", frac)
if o.onThinHigh != nil {
o.onThinHigh()
}
}
}
// NewObserver builds an Observer. host defaults to a ProcHostReader; logger to the
@@ -260,6 +300,7 @@ func (o *Observer) build(s proxmox.Storage, mounts []Mount) observed {
o.logger.Warn("storage: lvmthin pool data fill is high (a full pool corrupts every guest on it)",
"storage", s.Storage, "data_used_fraction", frac)
}
o.noteThinFill(s.Storage+"/data", frac)
}
// SMART-only device hint (v0.95.0): a dir-storage that lives INSIDE a shared filesystem (the
+34
View File
@@ -0,0 +1,34 @@
package storage
import (
"context"
"testing"
"gitea.dooplex.hu/admin/felhom-agent/internal/proxmox"
)
// R-672 (v0.133.0): a thin pool crossing 90 % requests an out-of-band host report ONCE, through the
// watchdog's own read path (Known — every few seconds), so the hub's storage-fill alarm sees the pool in
// seconds, not at the next 15-minute report. Re-armed below 85 %.
//
// COMPANION RED-PROOF (REPORT): remove the noteThinFill call from the data path → "no report was
// requested when the pool crossed 90 %".
func TestThinHigh_RequestsOneReportPerCrossing(t *testing.T) {
pool := proxmox.Storage{Storage: "local-lvm", Type: "lvmthin", Content: "rootdir,images", Total: 1000, Active: 1}
api := &fakeStorageAPI{node: "n", cluster: []proxmox.Storage{pool}}
o := NewObserver(api, &fakeHostReader{}, nil, quietLogger())
asked := 0
o.SetThinHighTrigger(func() { asked++ })
for i, used := range []int64{800, 910, 950, 1000, 840, 920} {
p := pool
p.Used, p.Avail, p.UsedFraction = used, 1000-used, float64(used)/1000
api.nodeSt = []proxmox.Storage{p}
if _, err := o.Known(context.Background()); err != nil {
t.Fatal(err)
}
want := map[int]int{0: 0, 1: 1, 2: 1, 3: 1, 4: 1, 5: 2}[i]
if asked != want {
t.Fatalf("after %d/1000 used: %d report requests, want %d", used, asked, want)
}
}
}
+12
View File
@@ -41,11 +41,23 @@ import sys
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
SHARED_REUSE = os.path.join(os.path.dirname(ROOT), "felhom.eu", "scripts", "reuse_refs_check.py")
SHARED_INSTRUCTIONS = os.path.join(
os.path.dirname(ROOT), "felhom.eu", "scripts", "instructions_gate.py")
# R-389 — shared, like the two above: it lives in felhom.eu/scripts/ and is never copied.
SHARED_OBSERVATIONS = os.path.join(
os.path.dirname(ROOT), "felhom.eu", "scripts", "observations_gate.py")
# (label, absolute script path, args, fast)
GATES = [
("reuse-refs", SHARED_REUSE, [ROOT], True),
("instructions", SHARED_INSTRUCTIONS, [ROOT], True),
("published", os.path.join(ROOT, "scripts", "check-published-versions.py"), [], False),
# R-273: the tag half of a release. Legs 1-2 need no network, so it runs in --fast too — the
# missing TAG is what actually broke every install, and the pre-push hook is the earliest place
# that can catch it.
("release-complete", os.path.join(ROOT, "scripts", "check-release-complete.py"), [], True),
# R-389 — a REPORT.md observation with no register row behind it. Fast: stdlib file reads.
("observations", SHARED_OBSERVATIONS, [ROOT], True),
]
VERDICT = {0: "OK", 1: "FAILED", 2: "INCONCLUSIVE"}
+41 -1
View File
@@ -95,6 +95,25 @@ PROBE_CONFIG = "configs/felhom-agent.service"
TAG_RE = re.compile(r"^v(\d+\.\d+\.\d+)$")
# THE retention number, read from the one file that owns it. A check and the policy it enforces
# must read the same number from the same place, or they drift and the drift looks like a defect
# in something else — which is exactly what happened on 2026-08-08/09 (R-287).
RETENTION_FILE = os.path.join(os.path.dirname(os.path.abspath(__file__)), "retention-policy.json")
def retention_kept():
"""How many of the newest generic versions the registry is expected to still serve.
Fails CLOSED and LOUD: a missing or unreadable policy file makes the check INCONCLUSIVE
rather than silently unbounded. An unbounded check would re-create the red this fixed; a
silently-bounded one would be worse.
"""
with open(RETENTION_FILE, encoding="utf-8") as fh:
n = json.load(fh)["generic_versions_kept"]
if not isinstance(n, int) or n < 1:
raise ValueError("generic_versions_kept must be a positive int, got %r" % (n,))
return n
tried = []
@@ -189,7 +208,28 @@ def main():
print(" no v<semver> tags in this repo yet — nothing to check, and nothing proven")
print("\ncheck-published-versions: NOTHING TO CHECK")
return 0
print(" %d released version(s) to verify: %s" % (len(versions), ", ".join(versions)))
all_versions = versions
try:
keep = retention_kept()
except Exception as e:
inconclusive("cannot read the retention policy (%s): %s" % (RETENTION_FILE, e))
# Bound the assertion to what the registry is expected to still hold. Sorted by SEMVER, not
# lexically: "0.9.0" > "0.10.0" as strings, and that would silently drop the wrong end.
def _key(v):
return tuple(int(x) for x in v.split("."))
versions = sorted(all_versions, key=_key)[-keep:]
dropped = [v for v in all_versions if v not in versions]
print(" %d released version(s); retention policy keeps the newest %d" % (len(all_versions), keep))
print(" verifying: %s" % ", ".join(versions))
if dropped:
# NEVER silent. A bounded check that does not say what it stopped covering is how a
# narrowing becomes permanent by accident.
print(" NOT ASSERTED (older than the retention window, and therefore not expected to be")
print(" downloadable): %s" % ", ".join(dropped))
print(" ^ these versions still have git TAGS and are still installable in the sense that")
print(" their configs resolve; what is no longer asserted is the BINARY's presence.")
bad = []
for v in versions:
+124
View File
@@ -0,0 +1,124 @@
#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""check-release-complete.py — the version at the head of CHANGELOG.md is a COMPLETE release.
THE DEFECT THIS IS A MACHINE FOR (2026-08-08/09, R-273). Agent v0.128.0 was built, tested,
CHANGELOG'd and published to the package registry — and its git tag was never pushed. The hub then
vouched it, and because felhom-host-install.sh fetches an agent's config files from
`raw/tag/v<version>/configs/`, EVERY fresh install and every reinstall died at step 5 of 8, as root,
on a virgin machine, for the better part of a day.
`scripts/release-agent.sh` already warns about exactly this, in as many words:
"a released version without a git tag 404s a box mid-install, as root"
The warning was there, it was correct, and the step was still missed. **So the fix is a machine and
not a reminder** — that is the whole point of this file.
WHAT IT ASSERTS, for the newest `## vX.Y.Z` in CHANGELOG.md:
1. a git tag `vX.Y.Z` EXISTS, and
2. it points at a commit that is an ANCESTOR OF (or equal to) the tip it was released from — a tag
parked on an unrelated commit is not a release, and
3. the generic package for X.Y.Z is DOWNLOADABLE.
(3) needs the network. (1) and (2) do not, and they are the half that actually failed — so this gate
is useful offline and says so rather than going quiet.
EXIT CODES, matching this repo's other gates: 0 clean, 1 convicted, 2 inconclusive. An unreachable
registry is INCONCLUSIVE for leg 3 only; legs 1 and 2 still run and can still convict.
"""
import json
import os
import re
import subprocess
import sys
import urllib.error
import urllib.request
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
GITEA_BASE = os.environ.get("GITEA_BASE", "https://gitea.dooplex.hu").rstrip("/")
OWNER, PKG = "admin", "felhom-agent"
HEAD_RE = re.compile(r"^##\s+v?(\d+\.\d+\.\d+)\b", re.M)
def git(*args):
return subprocess.run(("git",) + args, cwd=ROOT, capture_output=True, text=True)
def head_version():
ch = os.path.join(ROOT, "CHANGELOG.md")
if not os.path.exists(ch):
return None
m = HEAD_RE.search(open(ch, encoding="utf-8").read())
return m.group(1) if m else None
def main():
print("check-release-complete — the newest CHANGELOG version is a complete release")
v = head_version()
if not v:
print(" no '## vX.Y.Z' heading in CHANGELOG.md — nothing to check, and nothing proven")
return 0
tag = "v" + v
print(" newest CHANGELOG version: %s" % tag)
problems, inconclusive = [], []
# ---- leg 1 + 2: the tag, and where it points. Offline-capable. -----------------------------
r = git("rev-parse", "-q", "--verify", "refs/tags/%s^{commit}" % tag)
if r.returncode != 0:
# A shallow CI clone has no tags of its own; ask the remote before convicting, so this
# gate does not fire on a clone shape rather than on a real defect.
ls = git("ls-remote", "--tags", "origin", "refs/tags/%s" % tag)
if ls.returncode != 0:
inconclusive.append("cannot reach origin to look for tag %s: %s"
% (tag, ls.stderr.strip()[:120]))
elif not ls.stdout.strip():
problems.append(
"TAG %s DOES NOT EXIST. The installer fetches this version's configs from\n"
" %s/%s/felhom-agent/raw/tag/%s/configs/ — without the tag every install\n"
" 404s mid-run, as root. Fix: git tag -a %s <released-commit> && git push origin %s"
% (tag, GITEA_BASE, OWNER, tag, tag, tag))
else:
print(" ok tag %s exists on origin (not in this shallow clone)" % tag)
else:
sha = r.stdout.strip()
anc = git("merge-base", "--is-ancestor", sha, "HEAD")
if anc.returncode == 0:
print(" ok tag %s -> %s, an ancestor of HEAD" % (tag, sha[:10]))
else:
problems.append("tag %s points at %s, which is NOT an ancestor of HEAD — a tag parked "
"on an unrelated commit is not a release" % (tag, sha[:10]))
# ---- leg 3: the package. Needs the network. ------------------------------------------------
url = "%s/api/packages/%s/generic/%s/%s/%s" % (GITEA_BASE, OWNER, PKG, v, PKG)
req = urllib.request.Request(url, method="HEAD")
try:
with urllib.request.urlopen(req, timeout=25) as resp:
if resp.status == 200:
print(" ok package %s is downloadable" % v)
else:
problems.append("package %s returned HTTP %s at %s" % (v, resp.status, url))
except urllib.error.HTTPError as e:
if e.code == 404:
problems.append("PACKAGE %s IS NOT PUBLISHED (HTTP 404 at %s).\n"
" Fix: bash scripts/release-agent.sh %s" % (v, url, v))
else:
inconclusive.append("registry returned HTTP %s for %s" % (e.code, v))
except Exception as e:
inconclusive.append("registry unreachable (%s) — leg 3 not checked; legs 1-2 still ran" % e)
if problems:
print("\ncheck-release-complete: INCOMPLETE RELEASE")
for p in problems:
print(" - " + p)
return 1
if inconclusive:
print("\ncheck-release-complete: INCONCLUSIVE — an undetermined result is never a pass")
for i in inconclusive:
print(" - " + i)
return 2
print("\ncheck-release-complete: %s is tagged, placed and published." % tag)
return 0
if __name__ == "__main__":
sys.exit(main())
+44
View File
@@ -0,0 +1,44 @@
{
"_comment": [
"THE retention number for published agent artifacts. One file, read by everything that",
"depends on it, because a check and the policy it enforces must read the same number from the",
"same place or they drift — and the drift looks like a defect in something else.",
"",
"WHAT WENT WRONG WITHOUT IT (2026-08-08/09). The registry stopped serving felhom-agent",
"0.120.0 and older, while scripts/check-published-versions.py demanded that EVERY git tag",
"still be downloadable. Both rules are individually sensible; together they are impossible.",
"CI went red at a commit whose own run had been green the day before, on a true finding that",
"no one could act on. The red will return at the next publish unless the two read one number.",
"",
"HOW THE NUMBER WAS ARRIVED AT — stated honestly, because it is weaker than it looks.",
"generic_versions_kept is 10 because that is what the registry demonstrably holds today",
"(felhom-agent 0.121.0..0.128.0 = 10 versions, queried 2026-08-09). It is an OBSERVED state,",
"NOT a ruling anyone has been able to locate: no register row records a package prune, R-210",
"is WAITING-ON-OPERATOR and says 'Nothing was deleted; this is a list, not an action', and it",
"concerns local Docker images rather than this registry. Container packages currently hold 19",
"each, so there is no uniform ten-per-package cap visible either. See R-287.",
"",
"SO THIS FILE IS A FLOOR, NOT A LICENCE. It says: CI may assume nothing older than the newest",
"N generic versions is still downloadable. It does NOT authorise deleting anything, and the",
"operator should confirm or replace the number — at which point this file changes and both",
"readers follow it in the same commit.",
"",
"THE DEEPER BOUND, recorded so a future session does not have to re-derive it: the principled",
"limit is the hub's vouched min_agent floor (0.127.0 on 2026-08-09). Nothing can install an",
"agent below it — the hub refuses to vouch one and boxes update to the floor — so a released",
"version below the floor being un-downloadable costs nothing real. Bounding on the floor would",
"be better than bounding on a count, and it needs the gate to read the hub, which is network",
"the gate does not have today. Filed as the follow-up in R-287.",
"",
"NEVER retire a git TAG to satisfy this. felhom-host-install.sh fetches an agent's config",
"files from raw/tag/v<version>/configs/, so deleting a tag retires the ability to install that",
"version at all — a strictly worse act than an un-downloadable binary."
],
"generic_versions_kept": 10,
"readers": [
"scripts/check-published-versions.py — bounds its assertion to the newest N versions",
"documentation/runbooks/registry-retention.md (felhom.eu) — the prune procedure"
],
"recorded": "2026-08-09",
"recorded_by": "CC, from the registry's observed state; NOT from a located operator ruling"
}