Compare commits
5 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 6d7904786c | |||
| 856a127cd6 | |||
| 257c4d85c0 | |||
| 72161f6cf0 | |||
| 03b58cec0a |
+106
@@ -1,3 +1,109 @@
|
||||
## v0.125.0 — the agent opens the sealed bundle and returns one field (2026-08-04, R-199 links 7–8)
|
||||
|
||||
**Link 7 had one production caller and it was a `--selftest`.** `UnwrapIdentityBundle` has existed
|
||||
since slice 10D.1 and the only thing that ever called it was `runSelftestIdentityConsume`, reading the
|
||||
recovery code from an environment variable by hand. **Link 8 did not exist at all:** that selftest
|
||||
writes the whole bundle JSON to a file, and its success message named `tunnel_token + pbs_token` —
|
||||
an enumeration that was accurate when written and became a MISSTATEMENT the moment v0.77.0 sealed the
|
||||
offsite repository password into the same bundle. Anyone reading that output would conclude the
|
||||
repository password was not there. It now names what THIS bundle actually carried and what it did not.
|
||||
|
||||
**`POST /escrow/recover-offsite-password`** on the pinned local API: the controller supplies the
|
||||
customer's recovery code, the agent fetches this host's own sealed blob from the hub
|
||||
(`hub.Client.FetchIdentityEscrow`, hub ≥ v0.94.0, self-scoped by the per-host key), unseals it, and
|
||||
returns **only the offsite restic repository password** plus its sha256.
|
||||
|
||||
**Only that field, on purpose.** The bundle also carries the tunnel token, the PBS token and the WG
|
||||
private key. The controller is a trust tier down and needs none of them; returning them would widen
|
||||
the blast radius of a controller compromise for nothing. Narrowing costs nothing now and is not
|
||||
recoverable later.
|
||||
|
||||
**Why the agent and not the controller:** `age` is an agent runtime dependency and is deliberately
|
||||
absent from the controller image; the blob is a host-scoped object whose only writer is this agent
|
||||
under the per-host key, so the read is that write's mirror.
|
||||
|
||||
**R's handling is the tightest rule in this release.** It arrives in the request body over the pinned
|
||||
channel, is held in memory for one call, is cleared on the success path AND every failure path, is
|
||||
never written to disk, never an argument in a process list, never logged at any level including
|
||||
inside an error, and is never echoed. `UnwrapIdentity` already stages only the blob and the recovered
|
||||
plaintext in a temp dir it removes; a test redirects TMPDIR and asserts **the tree is empty
|
||||
afterwards** — emptiness rather than a content scan, because a content scan is defeated by a later
|
||||
call overwriting the leaked file, which is exactly how the first version of that test passed its own
|
||||
red-proof while R sat on disk.
|
||||
|
||||
Three outcomes are distinct rather than one generic failure: no blob (404 — no ceremony has run), a
|
||||
bundle that opens but predates the field (409 — a pre-fork-4 blob, which cannot be retro-fitted), and
|
||||
a code that does not open it (400 — fail-closed at the KDF, nothing written). Sending an operator to
|
||||
re-check a correctly typed recovery code because the hub was unreachable is the mistake this avoids.
|
||||
|
||||
**The wiring is asserted by an AST walk**, not a `strings.Contains`: `main` → `runDaemon` →
|
||||
`buildLocalAPIServer`, where an `escrow.OffsiteKeyRecoverer` is constructed and passed as
|
||||
`localapi.Options.EscrowRecovery`, and its fetcher calls the DAEMON's own hub client (the self-scoping
|
||||
that makes cross-host retrieval impossible is a property of which key is used). This project's
|
||||
built-but-never-wired count is six and links 6–7 were two of them; the fix must not become the seventh.
|
||||
|
||||
## v0.124.1 — the repair record must survive the probe that did NOT feed the hub (2026-08-04, R-190)
|
||||
|
||||
**v0.124.0's transition record did not reach the hub, and the live run is what showed it.** The
|
||||
capability reported degraded for "one cycle" — meaning the probe call that performed the repair. But
|
||||
`probeAll` is invoked **independently** by the periodic self-check log and by the collector building a
|
||||
host-report. On the demo box the repairing call was the log's (`09:39:34`, journal shows
|
||||
`GRANT WAS MISSING AND HAS BEEN SELF-REPAIRED` and `degraded=1`), and the host-report built three
|
||||
seconds later found the grant present and sent **`ok`**. The agent's journal had the record; the hub
|
||||
had nothing; the operator would have learned nothing.
|
||||
|
||||
That is the exact silence R-190 is about, re-created inside its own mitigation — and every unit test
|
||||
passed while it was true.
|
||||
|
||||
**The fix is a latch on TIME rather than on call count.** A confirmed repair is reported for
|
||||
`storeGrantRepairReportWindow` (20 minutes), which comfortably exceeds the 900 s host-report interval,
|
||||
so at least one report must carry the transition. It clears on its own — a permanently degraded
|
||||
capability would be its own false alarm — and it is per tier.
|
||||
|
||||
**Two hollow tests were caught and fixed on the way**, both the same shape this repo keeps finding: a
|
||||
test asserting a value it constructed itself, and a test asserting the latch HELPER rather than the
|
||||
path that consumes it — whose red-proof duly passed. The decisions now live in
|
||||
`storeGrantHealthyVerdict` and `storeGrantRepairedVerdict`, and the tests call those.
|
||||
|
||||
## v0.124.0 — a lost storage grant repairs itself, and says that it was lost (2026-08-04, R-190)
|
||||
|
||||
**R-190 is a grant that demonstrably worked at 04:44 on 2026-08-03 and was gone by 09:24** — with a
|
||||
host reinstall, logged `pveum` activity and cluster-log entries all ruled out by measurement. The
|
||||
cause is still open. The resilience does not have to wait for it.
|
||||
|
||||
**Everything needed already existed and had only ever been called once.** The root wrapper
|
||||
(`felhom-backup-target-apply grant <id>`), its sudoers vector (`grant *`, any storage id, confirmed
|
||||
not assumed), and the exact command were all in place — and the `grant` verb had only ever run at
|
||||
storage CREATION. That is the *built but never wired* shape, in a verb rather than a seam, and it is
|
||||
this project's seventh instance.
|
||||
|
||||
**What v0.124.0 does:** when the store-grant probe finds the grant absent on a tier the box depends
|
||||
on, it runs that wrapper and **re-reads once** to confirm — the pbsdr R-22 self-grant shape, including
|
||||
its restraint: one attempt, one confirmation, and anything still wrong stays loudly wrong.
|
||||
|
||||
**THE RECORD IS THE POINT, AND IT IS THE HALF R-190 IS ACTUALLY ABOUT.** A repair that leaves only
|
||||
`ok` behind destroys the only evidence a permission vanished, so a recurring loss becomes undetectable
|
||||
forever — strictly worse than the fault it fixes. So a confirmed repair reports **DEGRADED for exactly
|
||||
one cycle**, with the explanation in `Feature`:
|
||||
|
||||
```
|
||||
backup tier felhom-backup: the agent's storage grant was MISSING and has been AUTOMATICALLY
|
||||
RESTORED — the tier works now, but a permission that vanished on its own needs investigating (R-190)
|
||||
```
|
||||
|
||||
**Nothing new was built to carry it.** The hub's existing ok→degraded→ok edge is the channel — it
|
||||
alerts and e-mails on the first edge and logs the recovery on the next cycle, so one loss produces
|
||||
exactly one alert pair. No wire change, no hub change, no new event type. `Feature` carries the text
|
||||
because that is the field the hub interpolates into the operator's e-mail; `Reason` does not travel.
|
||||
|
||||
**Bounded (Scenario F):** one attempt per tier per hour, in memory. A storage can be unreadable for
|
||||
reasons an ACL cannot fix, and a re-grant on every report cycle is a repair loop wearing a fix's
|
||||
clothes. An agent restart re-arms it, which is correct — a restart is exactly when a box should
|
||||
re-check what it depends on.
|
||||
|
||||
**A failed repair never masks the fault:** the capability stays degraded with the failure in its
|
||||
reason, and a repair that "succeeded" but did not survive the re-read is reported as needing a human.
|
||||
|
||||
## v0.123.0 — a tier the box cannot READ now says so (2026-08-03, R-185)
|
||||
|
||||
**The missing permission is one command. The silence was the defect.** On demo-felhom the agent's PVE
|
||||
|
||||
+28
@@ -5,6 +5,34 @@
|
||||
|
||||
## Current
|
||||
|
||||
- **2026-08-03 — v0.123.0 (R-185): a tier the box cannot READ now says so.** The agent's token had
|
||||
`FelhomAgentStore` on `local`, `local-lvm`, `felhom-pbs` and **not** on `felhom-backup` — the
|
||||
storage both demo boxes configure as `local_backup_target`. That storage answered `{"data":[]}`
|
||||
through the token while root listed three archives, and `pickForThisRun` skipped it as *"no settled
|
||||
archive yet"* — **which is what a brand-new tier reports**, so the host tier was never
|
||||
restore-testable and nothing said so.
|
||||
- **The permission question is asked directly**, because unlike the listing it has a definite
|
||||
answer: `Client.Permissions` reads `/access/permissions?path=/storage/<target>` **as the agent's
|
||||
own token**, and `storeGrantStatuses` emits one `capability.Status` per configured tier. It
|
||||
composes AROUND the sudo prober, the way `poolReadStatus` already does — an API read does not
|
||||
belong inside a sudo-policy probe. `Status`'s wire shape is untouched, so the hub's critical
|
||||
degraded alert applies with **no hub change**.
|
||||
- **MEASURED FIRST, and the obvious reading is wrong:** an ungranted path answers neither empty nor
|
||||
403 — it carries the privileges INHERITED from the box-wide `/` grant
|
||||
(`Sys.Audit, SDN.Use, Datastore.Audit`). Checking path-presence, or `Datastore.Audit`, reports a
|
||||
blinded storage HEALTHY. The probe tests **`Datastore.AllocateSpace`**; re-measure before ever
|
||||
changing that constant (`storeGrantRequiredPriv`, red-proved).
|
||||
- **The probed set comes from `BackupTiers()`, never a fixed list** — a hardcoded probe list is the
|
||||
defect reproduced inside the fix. Critical, EXCEPT the `local` fallback target (reported, but it
|
||||
does not page). It never consults content, so it cannot alarm on a newborn tier; it never reports
|
||||
ok when it could not ask.
|
||||
- **LIVE:** degraded observed on the still-blind box (hub emailed `agent_capability_degraded`) →
|
||||
grant applied on **both** demo boxes → token lists 3 and 4 archives → `ok=70 total=70 degraded=0`
|
||||
and `degraded → ok` at the hub → **the host tier became a due-check candidate for the first time**,
|
||||
correctly picking the 08-02 archive (08-03 had not settled 24 h).
|
||||
- **The installer's real defect was NOT `PVE_STORAGES`** — see `felhom.eu` CONTEXT S-22: Case A
|
||||
grants, the Scenario-F reuse arm did not. Fixed in installer **1.24.0** with a gate.
|
||||
|
||||
- **2026-08-03 — v0.122.0 (R-189 · R-188 · R-186): three signals that lied about their own work.**
|
||||
None touches data; all three cost attention, which every other signal depends on.
|
||||
- **R-189 — a passing restore-test no longer vanishes on a restart.** `restore_tests[]` came only
|
||||
|
||||
@@ -1,9 +1,9 @@
|
||||
# REPORT — R-189 · R-188 · R-186: three ways the signals lied about themselves
|
||||
# REPORT — R-185: a tier the box cannot READ must say so
|
||||
|
||||
**Date:** 2026-08-03 · **Repo:** `felhom-agent` **v0.121.1 → v0.122.0** (`7581f81`) · released,
|
||||
published, verified by an independent download **and by rebuilding it**, deployed to demo-felhom.
|
||||
`felhom.eu`: register + docs only, **no hub change and no hub bump** — the hub already reads
|
||||
`restore_tests[]`; the defect was that the agent stopped sending them.
|
||||
**Date:** 2026-08-03 · **Repos:** `felhom-agent` **v0.122.0 → v0.123.0** (`fe14bc6`) · `felhom.eu`
|
||||
installer **1.23.0 → 1.24.0** (`688470c`, tag `installer-v1.24.0`, manifest bump `311dc06`) ·
|
||||
**no hub change and no hub bump** — the hub already alerts on a degraded critical capability, which is
|
||||
why that mechanism was chosen.
|
||||
|
||||
---
|
||||
|
||||
@@ -11,271 +11,195 @@ published, verified by an independent download **and by rebuilding it**, deploye
|
||||
|
||||
| Repo | `main` @ commit | Version | Matched §1? |
|
||||
|---|---|---|---|
|
||||
| `felhom-agent` | `3d0a1d615d11` | `v0.121.1` | **yes** |
|
||||
| `felhom.eu` | `c9a3e48b2106` | hub `v0.91.1` | **yes** |
|
||||
| `felhom-agent` | `0b28eae7bb14` | `v0.122.0` | **yes** |
|
||||
| `felhom.eu` | `7a5694341d59` | installer `1.23.0`, both `--ref=installer-v1.23.0` (lines 327, 372) | **yes** |
|
||||
|
||||
Highest register ID in use was **R-189**; no new IDs were needed — all three rows already existed.
|
||||
(Grep confirmed R-190+ free, in case one had been.)
|
||||
Highest register ID in use **R-189**; R-190+ confirmed free by grep, and none was needed.
|
||||
|
||||
## 2. Scenario H — the reproducibility measurement (Part 3, done first on purpose)
|
||||
## 2. Part 0 — the measurements, before anything was designed against them
|
||||
|
||||
**Before**, one commit, same source, same toolchain, same ldflags — the only difference is whether the
|
||||
tag existed when the build ran:
|
||||
**The row's three-way observation, reproduced unchanged:**
|
||||
|
||||
| build | sha256 | size | embedded module version |
|
||||
|---|---|---|---|
|
||||
| default flags, **no tag yet** | `18f4a495…` | 14 085 464 B | `v0.121.2-0.20260803133646-3d0a1d61` |
|
||||
| default flags, **tagged** | `4a38f394…` | 14 085 440 B | `v0.121.99` |
|
||||
| `-trimpath -buildvcs=false`, either way | `7ffcdf1d…` | 14 064 574 B | *(none)* |
|
||||
| leg | result |
|
||||
|---|---|
|
||||
| root, `pvesh … /storage/felhom-backup/content` | **3 archives** — 6.1 / 6.2 / 6.3 GB, dated 08-01, 08-02, 08-03 |
|
||||
| the **agent's token**, same endpoint | `{"data":[]}` |
|
||||
| the agent's token, `/storage/local/content` | **8 entries** — the token works where it is granted |
|
||||
|
||||
**After, on the real release (v0.122.0) — the three values §15.2 asks for:**
|
||||
So the token is the variable, not the storage. Two further checks removed the obvious alternative
|
||||
explanation: guest **9201 IS in the `felhom` pool** (so `VM.Backup` is not the discriminator), and
|
||||
`pveum acl list` showed ACL rows for `/storage/{local,local-lvm,felhom-pbs}` and **none** for
|
||||
`/storage/felhom-backup`.
|
||||
|
||||
| artifact | sha256 | size |
|
||||
|---|---|---|
|
||||
| published, downloaded from Gitea | `d5f294e56c1ef59055e8e87fb9135aa477632dbbc56a5d4bff46bbd0466c1edf` | 14 076 649 B |
|
||||
| rebuild at the tag, #1 | `d5f294e56c1ef59055e8e87fb9135aa477632dbbc56a5d4bff46bbd0466c1edf` | 14 076 649 B |
|
||||
| rebuild at the tag, #2 | `d5f294e56c1ef59055e8e87fb9135aa477632dbbc56a5d4bff46bbd0466c1edf` | 14 076 649 B |
|
||||
|
||||
**All three identical.** The property is removed-cause, not sequenced-around: `-buildvcs=false` drops
|
||||
a stamp nothing reads (no `ReadBuildInfo` caller, verified by grep), and the version still comes from
|
||||
the explicit `-X main.version` ldflag. `-trimpath` additionally makes a rebuild from a different
|
||||
checkout directory match.
|
||||
|
||||
**A second discrepancy fell out of the measurement and is fixed with it:** `publish-agent.sh`'s
|
||||
fallback build forced `CGO_ENABLED=0` and produced **13 990 236 B** against the release path's
|
||||
**14 064 574 B** — a 74 KB difference, i.e. one version name meaning two binaries depending on which
|
||||
entry point ran. Both paths now use identical flags, each commented with a pointer to the other.
|
||||
|
||||
## 3. Scenario E — a correct release no longer emails a failure (Part 2)
|
||||
|
||||
**Only the tag PUSH moved.** The order is now build → tag **locally** → publish → push tag. The tag is
|
||||
still created before anything is published, so the build and the tag describe the same commit; it
|
||||
becomes *visible* — to CI (`on: [push]`) and to any `raw/tag/…` fetch — only once the package is
|
||||
downloadable.
|
||||
|
||||
**The old ordering's invariant is asserted directly rather than arranged for.**
|
||||
`check-published-versions.py` now carries two invariants: every tag has an installable package (as
|
||||
before) **and no published version is missing its tag** (new). The second is a **bounded probe** —
|
||||
the frontier past the newest tag, where a failed tag push leaves an orphan, plus patch gaps — and it
|
||||
**prints its probe set on every run**, because a check whose coverage is invisible reads as a
|
||||
guarantee it is not making. The package listing api was **re-measured**, not assumed: `401` without a
|
||||
token, so absence still cannot be enumerated, and the script says so in its own output.
|
||||
|
||||
**Live result — this release is the test:**
|
||||
|
||||
| release | CI runs on the release commit | outcome |
|
||||
|---|---|---|
|
||||
| v0.121.0 (yesterday) | #12 / #13 | one **green**, one **red** |
|
||||
| v0.121.1 (yesterday) | #17 / #18 | one **red**, one **green** |
|
||||
| **v0.122.0 (this one)** | **#21 (task id 96) / #22 (task id 97)** | **both green** |
|
||||
|
||||
**Failure modes made loud rather than tidy:** a publish that succeeds followed by a tag push that
|
||||
fails now dies printing `git push origin v<ver>` — the local tag is already there, so recovery is one
|
||||
line — and a publish that *fails* deletes the local-only tag so a retry is clean instead of colliding
|
||||
with step 2's re-release guard.
|
||||
|
||||
## 4. Scenarios F and G — both directions demonstrated, then cleaned up
|
||||
|
||||
**G — a published version with no tag must FAIL.** A real fixture: version **0.121.2** (the frontier,
|
||||
exactly where a failed tag push lands) published to the live registry with no tag.
|
||||
**The permission query, asked by the token itself — and the obvious reading is wrong:**
|
||||
|
||||
```
|
||||
0.121.2 next patch after the newest tag PUBLISHED — NO TAG
|
||||
check-published-versions: 1 PUBLISHED VERSION(S) WITH NO TAG
|
||||
v0.121.2 is downloadable at … but has no git tag.
|
||||
git push origin v0.121.2
|
||||
EXIT=1
|
||||
/storage/felhom-pbs → {"Datastore.Allocate":1,"Datastore.AllocateSpace":1}
|
||||
/storage/felhom-backup → {"Sys.Audit":1,"SDN.Use":1,"Datastore.Audit":1}
|
||||
```
|
||||
|
||||
**Red-proof (observed):** with the fixture still live, replacing the probe set with `[]` produced
|
||||
`ALL RELEASED VERSIONS INSTALLABLE, AND NONE UNTAGGED`, **exit 0** — a green run over a published
|
||||
orphan. Restored.
|
||||
The ungranted path answers **neither empty nor 403**. It answers with the privileges **inherited**
|
||||
from the box-wide `/` grant. A probe asking *"did the path come back?"* — or *"does it hold
|
||||
`Datastore.Audit`?"* — would have reported the blinded storage **healthy**. This is exactly what §3
|
||||
required to be measured rather than assumed, and it changed the design: the probe tests
|
||||
`Datastore.AllocateSpace` specifically, and a red-proof pins that choice.
|
||||
|
||||
**Teardown:** fixture deleted (HTTP 204), absence independently re-verified (`GET → 404`), gate green
|
||||
again (exit 0). No scratch tag was ever pushed.
|
||||
## 3. The probe
|
||||
|
||||
**F — a tag with no package must still FAIL.** Demonstrated against a local stand-in serving two tags
|
||||
where only one has a package:
|
||||
`Client.Permissions` reads `/access/permissions?path=/storage/<target>` **as the agent's own token**
|
||||
(asking as root answers a different question and always says yes). `storeGrantStatuses` emits one
|
||||
`capability.Status` per configured tier.
|
||||
|
||||
```
|
||||
ok v0.120.0: binary downloadable + tag serves its configs
|
||||
FAIL v0.199.0:
|
||||
- binary NOT downloadable (HTTP 404 …)
|
||||
- tag does not serve configs/felhom-agent.service (HTTP 404) — a box would 404 mid-install
|
||||
EXIT=1
|
||||
```
|
||||
**Deviation from §5/§8.1, stated because a recommendation not followed gets a line:** the spec asked
|
||||
for the sudo `Prober` to be minimally generalised. This repo already has the better-established
|
||||
pattern for exactly this — `poolReadStatus`, composed **around** the prober, with the comment *"an API
|
||||
read does not belong inside the sudo-policy probe"* (v0.62.0, audit A1). The probe follows that
|
||||
precedent instead. `capability.Status` is untouched either way, which is the constraint that mattered.
|
||||
|
||||
**Why not a real pushed tag:** pushing one wakes CI and would have emailed the operator a **true**
|
||||
alarm about a fixture — the same attention cost R-188 exists to remove. F is also already
|
||||
demonstrated in the wild: CI runs **#13** and **#17** failed for exactly this reason yesterday.
|
||||
**Decisions:**
|
||||
|
||||
## 5. R-189 — a proof that survives a restart reaches the hub (Part 1)
|
||||
- **The probed set comes from the box's own `BackupTiers()`**, never a fixed list — a hardcoded probe
|
||||
list is the defect reproduced inside the fix.
|
||||
- **Critical** (§8.3): the hub alerts only on critical, so a non-critical entry would ride the report
|
||||
and alert nobody — the same silence with extra steps. **Except** the `local` fallback target, which
|
||||
host-install's own comment calls the DEGRADED configuration: still probed, still reported, but it
|
||||
does not page. Turning an ordinary documented setup into an alert is how a signal becomes something
|
||||
an operator archives unread.
|
||||
- **It never consults content**, so it cannot alarm on a newborn tier by construction — a stronger
|
||||
guarantee than gating on emptiness would be.
|
||||
- **It never reports ok when it could not ask.** Unreachable PVE is degraded: a self-check that fails
|
||||
open converts *"I do not know"* into *"fine"*.
|
||||
|
||||
`restore_tests[]` came only from the in-memory `backup.Store`, whose comment read *"lost on restart;
|
||||
the cadence re-populates"*. True under a timer; false since R-86, because the agent refuses to re-test
|
||||
an archive it has already proven — so a lost proof is not repeated for a whole archive generation.
|
||||
## 4. The installer — the root cause was not where the row or the task expected
|
||||
|
||||
**What changed**
|
||||
Both assumed `PVE_STORAGES` (the fixed grant list) was the culprit. **It is not.**
|
||||
`configure_backup_target` has two arms:
|
||||
|
||||
- `RestoreTestState` stores the **tier** and what was **verified** beside the archive (v3 shape).
|
||||
Both are recorded **at proof time from the run's own result** — deriving them later would need a
|
||||
storage-type lookup at report-building time, a network call that can fail on the one path where
|
||||
failing means mislabelling a proof.
|
||||
- `ProvenRestoreTests` renders the stored proofs as report entries; `Collector.SetProvenRestoreTests`
|
||||
merges them with the in-memory result.
|
||||
- **Merge rule: one entry per tier, newest by `TestedAt` wins.** It falls out of what each source
|
||||
means rather than from a preference: a fresh failure beats a stored success (the failure is the
|
||||
news and lives nowhere else), a stored success beats a stale in-memory entry after a restart, and a
|
||||
tier never appears twice — the hub would read that as two tests. An unparseable timestamp counts as
|
||||
**older**, so a malformed entry cannot displace a good one.
|
||||
- **It refuses to lie.** A record missing the archive **or** the tier produces no entry, and run
|
||||
mechanics (scratch VMID, duration) are not re-invented: an absent duration is not a claim, a
|
||||
fabricated one would be.
|
||||
- **The asymmetry is now in the code** (§8.1): a success *suppresses* future work so it must be
|
||||
durable; a failure *causes* future work and heals itself, and persisting one would make a healed
|
||||
tier keep reporting a fault.
|
||||
- **`Store`'s comment is corrected in place** — leaving it is how the next reader concludes this is
|
||||
handled.
|
||||
- **Case A** creates the storage and calls `felhom-backup-target-apply grant` in the same breath — a
|
||||
box that builds its own target has always been correct.
|
||||
- **The Scenario-F arm** — *"the target already exists, leave it exactly as it is"* — **returned
|
||||
without granting**.
|
||||
|
||||
**Migration, and it is visible on the live box:** a pre-R-189 record has an archive but no tier, so it
|
||||
is **not** reportable. Upgrading does not retroactively make an old proof visible; the tier's next
|
||||
real proof fills it in. Confirmed immediately after the deploy — still `0 restore-tests`, with the
|
||||
v2 record sitting on disk.
|
||||
So a box whose `felhom-backup` pre-dated the install (created by the vzdump-target-move runbook, or
|
||||
surviving a reinstall — which is both demo boxes) pointed `local_backup_target` at a storage its own
|
||||
token could not read. The reuse arm now ensures the ACL through the same guarded wrapper.
|
||||
|
||||
## 6. Scenario A, live on demo-felhom — against the observation that filed R-189
|
||||
**Scenario F is unviolated:** the storage DEFINITION is still untouched. Granting the role the agent is
|
||||
supposed to have on the target this same script is about to write into `agent.json` is finishing the
|
||||
job, not retargeting the box; `pveum acl modify` is idempotent, so a box that already has it is
|
||||
unchanged and a box whose token was rotated gets it back.
|
||||
|
||||
**The observation being replaced (2026-08-03, 15:25):** a real offsite restore-test PASSED, the agent
|
||||
was restarted 2 m 43 s later, and the hub logged `0 restore-tests` on the next two host-reports.
|
||||
**`$BACKUP_TARGET_ID` is deliberately still NOT in `PVE_STORAGES`,** and the comment now says why: that
|
||||
list is granted in step 4/5, *before* `configure_backup_target` runs in step 6, and `--acl-storages`
|
||||
entries are preflight-checked for existence. Adding it there would grant on a storage that may not yet
|
||||
exist and would split ownership of the decision across two places.
|
||||
|
||||
**The same sequence, on v0.122.0:**
|
||||
**A gate now asserts it:** every arm of `configure_backup_target` that resolves the target must also
|
||||
grant on it — the check that would have caught this.
|
||||
|
||||
```
|
||||
16:44:06 restore-test tier is DUE target=felhom-pbs
|
||||
archive=felhom-pbs:backup/ct/9201/2026-07-27T19:55:41Z
|
||||
reason="newest settled archive … has not been proven (last proven archive was a different one)"
|
||||
16:55:21 restore-test: scratch guest torn down vmid=990000
|
||||
16:55:21 backup: scheduled restore-test PASSED archive=felhom-pbs:…2026-07-27T19:55:41Z duration_s=675.1
|
||||
16:55:32 systemctl restart felhom-agent ← INSIDE the 15-minute reporting window
|
||||
16:55:36 hub: host-report from demo-felhom-8363b5 (… 1 restore-tests …) ← was 0
|
||||
```
|
||||
## 5. Live validation, in order
|
||||
|
||||
**The proof on disk (v3 — the tier is what the old shape lacked):**
|
||||
| # | evidence |
|
||||
|---|---|
|
||||
| 1 | Part 0's measurements above, taken **before** any change |
|
||||
| 2 | **The signal that has never existed**, on the still-blind box: `capability DEGRADED … capability=pve:store-grant:felhom-backup … reason="the agent token lacks Datastore.AllocateSpace on /storage/felhom-backup (grant FelhomAgentStore there) — this tier's archives are INVISIBLE to the agent and it is never restore-tested" critical=true`, with `ok=69 total=70 degraded=1`. The hub: `Host capability: demo-felhom-8363b5 ok → degraded (agent_capability_degraded)` and **`Operator email sent`** |
|
||||
| 3 | Grant applied (user **and** token — a privsep token's rights are the intersection); the token then lists **3 archives** where it listed none, and the permission answer becomes `{"Datastore.AllocateSpace":1,"Datastore.Allocate":1}` |
|
||||
| 4 | `capabilities self-check ok=70 total=70 degraded=0`; the hub: `degraded → ok (agent_capability_recovered)` |
|
||||
| 5 | **The host tier is a due-check candidate for the first time on that box**: `tier=felhom-backup due=true archive="…2026_08_02-04_42_14.tar.zst" proven=""` — and the settle rule applies to it exactly as to the others, selecting the **08-02** archive because the 08-03 one has not settled 24 h |
|
||||
| 6 | The served installer over HTTPS: `SCRIPT_VERSION="1.24.0"`, and the served bytes carry the fix itself, not merely the version |
|
||||
|
||||
```json
|
||||
{"felhom-pbs": {"archive": "felhom-pbs:backup/ct/9201/2026-07-27T19:55:41Z",
|
||||
"tier": "pbs", "verified": "boot+running", "proven_at": "2026-08-03T14:55:21Z"}}
|
||||
```
|
||||
|
||||
**What the HUB stored** — read from its own database (copied with its `-wal`, freshness confirmed by
|
||||
the newest row's `received_at` = `2026-08-03 14:55:36` UTC, matching the ingest line):
|
||||
|
||||
```json
|
||||
{ "source_archive": "felhom-pbs:backup/ct/9201/2026-07-27T19:55:41Z",
|
||||
"source_tier": "pbs", "pass": true, "verified": "boot+running",
|
||||
"tested_at": "2026-08-03T14:55:21Z", "scratch_vmid": 0, "duration_seconds": 0 }
|
||||
```
|
||||
|
||||
That report was built **after** the restart, when the in-memory store was empty — so the entry can
|
||||
only have come from the persisted state. The archive, the tier, and the **original** test time
|
||||
survived; the run mechanics are zero because they are deliberately not re-invented.
|
||||
|
||||
**A 14.5 GB encrypted offsite archive**, restored, booted, verified and destroyed in **675 s** — and
|
||||
this time the proof outlived the process that produced it.
|
||||
|
||||
**Teardown, all three layers:** scratch guest absent from `pct list` (0), its volumes gone from `lvs`
|
||||
(0), the validation drop-in removed and the daemon back on its defaults
|
||||
(`eval_interval=6h0m0s settle=24h0m0s`). The hub-side `restore_tests[]` record is **retained
|
||||
deliberately** — it is the proof the staleness check reads, so deleting it would delete the result.
|
||||
No `restore_test_*` event was raised, because nothing failed and nothing is stale.
|
||||
## 6. The other machines
|
||||
|
||||
- **demo-hp CARRIES THE SAME DRIFT — and was fixed.** `local_backup_target=felhom-backup`, ACL rows for
|
||||
`local`, `local-lvm`, `felhom-pbs` only. §8.6 assumed a single affected box; the same one-line,
|
||||
additive, path-scoped, idempotent grant applies to the other, and leaving a known-blind backup tier
|
||||
on a Tier-0 box after finding it would be this row happening twice. Granted (user + token); its
|
||||
token now lists **4 archives**. It still runs agent `0.120.0`, so it has no probe yet — that arrives
|
||||
when you vouch.
|
||||
- **The tester's box was NOT touched** (Tier 2, protected). **What is known without connecting to it:**
|
||||
it very likely carries the same drift — the mechanism is the Scenario-F reuse arm, which fires on
|
||||
any box whose target pre-dated its install, and its target was moved by the very runbook that
|
||||
creates that condition. It is due for reinstall, and installer 1.24.0 fixes it on the way in.
|
||||
|
||||
## 7. Tests and red-proofs
|
||||
|
||||
Green gate: `go build ./... && go vet ./... && go test ./...` — **29 packages ok, rc=0**, plus
|
||||
`python3 scripts/agent_gates.py` (reuse-refs + published-versions) all OK. The test run and the commit
|
||||
were always separate commands.
|
||||
Green gate: `go build ./... && go vet ./... && go test ./...` — rc=0, plus `agent_gates.py` and
|
||||
`repo_gates.py` all OK. Test runs and commits were always separate commands.
|
||||
|
||||
| # | Test | Asserts | Mutation | Observed |
|
||||
|---|---|---|---|---|
|
||||
| A | `TestMerge_ProofSurvivesARestart` | an empty in-memory store + a persisted proof → the proof is reported, with its archive and its original time | the persisted merge deleted (the pre-R-189 body) | **FAIL** — `after a restart the persisted proof must be reported; got 0 entr(ies): []` — the live observation exactly |
|
||||
| B | `TestMerge_NeverInventsAPassForAnUnprovenTier` | no proof → no entry; a tier-less record → no entry | — (its state-layer twin below carries the mutation) | pass |
|
||||
| B′ | `TestProvenRestoreTests_RefusesToReportWhatItCannotDescribe` | v1 + v2 + v3 records side by side → only the describable one is reported | the `reportable()` filter dropped | **FAIL** — `got 3` entries, two with an empty `SourceTier`/`SourceArchive` |
|
||||
| C | `TestMerge_NewerWinsAndNeverDuplicatesATier` | one entry per tier, newest wins, in both directions | de-duplication removed | **FAIL** — `one entry per tier; got 2 for "pbs" — the hub would read two tests` |
|
||||
| D | `TestMerge_AFailureIsStillReported` | a fresh failure beats an older stored success | (same mutation) | **FAIL** — 2 entries, i.e. the failure no longer the single answer for that tier |
|
||||
| — | `TestMerge_MalformedTimestampNeverWins` | unparseable ≠ newest | — | pass |
|
||||
| — | `TestMerge_NilProvenSourceIsANoOp` | pre-R-189 behaviour unchanged when unwired | — | pass |
|
||||
| — | `TestScheduler_ProofIsRecordedReportably` | a pass **through the scheduler** leaves a reportable proof | — | pass |
|
||||
| — | `TestScheduler_AFailureLeavesNoPersistedProof` | §8.1's asymmetry, asserted not assumed | — | pass |
|
||||
| G | the gate's converse assertion | a published version with no tag fails | probe set → `[]` | **FAIL** (green over a live orphan) |
|
||||
| I | `TestMainWiresTheDurableRestoreTestProof` | **AST**: `SetProvenRestoreTests` is called **and fed `rtState`** | the call commented out | **FAIL** — `main.go never calls collector.SetProvenRestoreTests` (a `strings.Contains` check would have passed — the string is still there) |
|
||||
| H | reproducibility | three identical sha256 | — (measurement, §2) | pass |
|
||||
| A | `TestStoreGrant_ForbiddenStorageIsDegradedAndNamed` | degraded, critical, naming storage **and** role | probe removed from `probeAll` | **FAIL** — `main.go never calls storeGrantStatuses` (via the seam test); with the wrong-privilege mutation: `must be DEGRADED, not "ok"` |
|
||||
| A′ | `TestStoreGrant_InheritedPrivilegesAreNotAGrant` | the measured trap: inherited ≠ granted | probe `Datastore.Audit` instead | **FAIL** — `checking for the wrong privilege reports a blinded storage healthy; got "ok"` |
|
||||
| B | `TestStoreGrant_GrantedButEmptyIsHealthy` | a readable-but-empty tier is healthy | — (it never reads content, so emptiness cannot reach it) | pass |
|
||||
| B′ | `TestStoreGrant_TheFallbackTargetIsNotCritical` | `local` is reported but does not page | gating removed (`return true`) | **FAIL** — `must not page the operator about an ordinary, documented configuration` |
|
||||
| C | `TestStoreGrant_ForbiddenAndNewbornAreDistinguishable` | different status **and** different capability id | — | pass |
|
||||
| — | `TestStoreGrant_UnreachablePVEIsDegradedNotOK` | unknown ≠ ok | — | pass |
|
||||
| F | `hostinstall_gates.py` backup-target assertion | every resolving arm also grants | reuse arm reverted | **FAIL** — `resolves the backup target in 2 place(s) but grants in only 1` |
|
||||
| H | `TestMainWiresTheStoreGrantProbe` | **AST** of `main.go` | call commented out | **FAIL** — a `strings.Contains` check would have passed |
|
||||
|
||||
**Jitter, per §10:** every timestamp fixture uses odd minutes and seconds (`13:25:14`, `19:55:41`,
|
||||
`13:41:07`, `04:41:58`) — several taken from the real box — rather than round hours. Yesterday a test
|
||||
was hollow because a perfectly regular series landed exactly on a threshold and survived its own
|
||||
mutation.
|
||||
**A hollow test caught and fixed before it shipped:** the first draft of `storegrant_test.go`
|
||||
re-implemented the verdict branch inside the test. It passed, and would have kept passing while
|
||||
production diverged. The decision was extracted into `storeGrantVerdict` and the tests now call it.
|
||||
|
||||
## 8. Files changed
|
||||
**Scenario B's red-proof, honestly:** the spec asked for "degrade on an empty content listing" as the
|
||||
mutation. That is not a mutation of this code — the probe never looks at content, which is a stronger
|
||||
guarantee than gating on emptiness. The gating red-proof above (`storeGrantCritical`) is the one that
|
||||
exercises the guard that does exist, and it fails as required.
|
||||
|
||||
`internal/backup/restoretest_state.go` (v3 record + `ProvenRestoreTests`), `internal/backup/store.go`
|
||||
(the comment that had become false), `internal/backup/schedule.go` (record tier + verified),
|
||||
`internal/hub/collect.go` (the seam + the merge), `cmd/felhom-agent/main.go` (wiring),
|
||||
`scripts/release-agent.sh` (ordering, reproducible build, loud half-done release),
|
||||
`scripts/publish-agent.sh` (identical build flags), `scripts/check-published-versions.py` (the
|
||||
converse invariant), plus `REUSE.md`, `CHANGELOG.md`, `CONTEXT.md`, `CLAUDE.md` and three test files.
|
||||
## 8. Files, commits, tag
|
||||
|
||||
**Commits** — `felhom-agent`: `7581f81` (v0.122.0). `felhom.eu`: see §10.
|
||||
`internal/proxmox/query.go` (`Permissions`), `cmd/felhom-agent/main.go` (`storeGrantStatuses`,
|
||||
`storeGrantVerdict`, `storeGrantCritical`, `storeGrantRequiredPriv`, wiring),
|
||||
`cmd/felhom-agent/storegrant_test.go`, `CHANGELOG.md`, `CONTEXT.md`, `REUSE.md`, `REPORT.md`.
|
||||
`felhom.eu`: `scripts/felhom-host-install.sh`, `scripts/hostinstall_gates.py`, `scripts/CHANGELOG.md`,
|
||||
`manifests/webpage.yaml`, `CONTEXT.md`, `STATUS.md`, `documentation/architecture/00-capability-map.md`,
|
||||
`documentation/backlog/OPEN-ITEMS.md`, `documentation/runbooks/RUNBOOK-vzdump-target-move-2026-07-29.md`.
|
||||
|
||||
## 9. The independent-verification command (Part 3, recorded in `CLAUDE.md`)
|
||||
**Commits** — `felhom-agent`: `fe14bc6` (v0.123.0). `felhom.eu`: `688470c` (installer 1.24.0), `311dc06`
|
||||
(manifest refs), `e3187c8` (docs). **Installer tag:** `installer-v1.24.0`.
|
||||
|
||||
```bash
|
||||
V=0.122.0
|
||||
git checkout "v$V" && go build -trimpath -buildvcs=false -ldflags "-X main.version=$V" \
|
||||
-o /tmp/felhom-agent-check ./cmd/felhom-agent
|
||||
sha256sum /tmp/felhom-agent-check
|
||||
curl -fsSL "https://gitea.dooplex.hu/api/packages/admin/generic/felhom-agent/$V/felhom-agent" | sha256sum
|
||||
```
|
||||
## 9. Deployment
|
||||
|
||||
Both print `d5f294e56c1ef59055e8e87fb9135aa477632dbbc56a5d4bff46bbd0466c1edf` (§2).
|
||||
Agent released through `release-agent.sh` — tag `v0.123.0`, sha256
|
||||
`74910135ac4feb1b7f0ad4dbd1541d965cbc0fe70d4f47b62ebf7e4bfb962453`, round-trip verified. The
|
||||
**published bytes** were downloaded and deployed: the running binary's sha matches the published one.
|
||||
`felhom-agent --version` → **0.123.0**, `systemctl is-active` → active, prior kept as `.bak-0.122.0`.
|
||||
**NOT VOUCHED** — that stays the operator's act.
|
||||
|
||||
## 10. Registers
|
||||
|
||||
- **R-189 → CLOSED** (shipped + proven live), **R-188 → CLOSED** (shipped), **R-186 → CLOSED**
|
||||
(shipped + measured).
|
||||
- **R-185 remains OPEN and untouched** — it is a missing `/storage/felhom-backup` ACL on demo-felhom,
|
||||
a permission defect, not a reporting one. Nothing in this session changed it, and the priority list
|
||||
says so explicitly.
|
||||
- No new IDs minted. `ROADMAP.md` contains none of these three rows, so there was nothing to collapse.
|
||||
- `00-capability-map.md`'s restore-proof row now records that the evidence path itself had a gap and
|
||||
what closed it; `CONTEXT.md` gains **S-19** (the proof/failure asymmetry and the merge rule) and
|
||||
**S-20** (the release ordering and what each step protects); `STATUS.md` rewritten for the operator
|
||||
and trimmed to 83 lines.
|
||||
- **R-185 → CLOSED** (shipped + proven live on both demo boxes), with the corrected root cause
|
||||
recorded on the row.
|
||||
- No new IDs minted; `ROADMAP.md` contains no R-185 row, so there was nothing to collapse.
|
||||
- **The capability map's whole-guest row was OPTIMISTIC and now says so:** every live restore-test it
|
||||
cited is on the OFFSITE tier, and the HOST tier was not merely unproven but *unprovable* on both
|
||||
demo boxes. It now records that, the closure, and that it will carry a host-tier live proof when one
|
||||
runs.
|
||||
- The vzdump-target-move runbook's item 5 **predicted this** and is annotated, not rewritten: it
|
||||
expected a 403 on backup, and the reason it did not surface that way is that `vzdump` writes through
|
||||
a root path, so backups kept landing while the agent's *read* stayed blind.
|
||||
- `CONTEXT.md`: agent-side entry, plus `felhom.eu` **S-21** (empty ≠ forbidden; the measured trap) and
|
||||
**S-22** (the Scenario-F arm must finish the job).
|
||||
|
||||
## 11. Observations — noticed, recorded, NOT acted on
|
||||
## 11. Teardown
|
||||
|
||||
- **The proof state holds ONE record per tier, so proving an OLDER archive re-arms a newer one —
|
||||
CONFIRMED after the validation, not merely predicted.** With the defaults restored, the due-check
|
||||
reads: `tier=felhom-pbs due=true archive="…2026-07-28T04:49:43Z" proven="…2026-07-27T19:55:41Z" —
|
||||
newest settled archive has not been proven (last proven archive was a different one)`.
|
||||
Surfaced by this session's own validation method: to get a fresh proof without waiting a week, the
|
||||
settle lag was widened so the older, unproven offsite archive became the candidate. That overwrote
|
||||
the record for the newer archive, so once the default 24 h settle returns, the newest settled
|
||||
archive is no longer the recorded proof and the tier becomes due once more. **Consequence, stated
|
||||
rather than left to surprise: demo-felhom will run one further unattended offsite restore-test
|
||||
within 6 h, after which the newest archive is the recorded proof** — the correct steady state. In
|
||||
normal operation this cannot arise, because the candidate only ever moves forward.
|
||||
- **`RestoreTestState.Snapshot()` has no caller again.** The host report is now fed by
|
||||
`ProvenRestoreTests`, which carries what a bare timestamp cannot. The method's doc comment says in
|
||||
as many words that it should be deleted if it does not acquire one — deliberately not deleted in
|
||||
this session, because removing an exported method is a change with no bearing on the three rows.
|
||||
- **Ten files in this repo are not `gofmt`-clean and were already so on arrival**
|
||||
(`internal/capability/probe.go`, `internal/escrow/consume.go`, `internal/mgmtplane/mgmtplane.go`,
|
||||
`internal/reconcile/bringup.go`, `internal/signedjobs/runner.go`, `internal/storage/{candidates,
|
||||
intent}.go` and three test files). Every file this session touched is clean; the others are
|
||||
untouched, and no gate checks formatting.
|
||||
- **The hub sweeps every 60 s and re-reads 14 days of host-reports per customer** for the restore-test
|
||||
staleness check. Unchanged here and not a defect at this fleet size; it is the cost centre if the
|
||||
fleet grows, and it is the reason the window read was deliberately left at 14 days yesterday.
|
||||
- **`felhom.eu/CONTEXT.md` still carries duplicate standing-ruling IDs** (three `S-14`s, two `S-15`s)
|
||||
from before yesterday. New rulings continue to be numbered above the collision (S-19, S-20) rather
|
||||
than adding to it; renumbering the existing ones is a separate, purely editorial change.
|
||||
**Nothing was provisioned.** No scratch storage, no fixture grant, no probe tag, no scratch package
|
||||
version. The two ACL grants are the intended durable change; the only other mutation was the
|
||||
installer label, which is reversible by moving the tag.
|
||||
|
||||
## 12. Observations — noticed, recorded, NOT acted on
|
||||
|
||||
- **Both demo boxes are now due for a host-tier restore-test**, which has never run on either. The
|
||||
scheduler will pick it up within 6 h unattended (a ~6 GB local restore — fast, and cheaper than the
|
||||
offsite ones). Expected, not a defect, and the first host-tier proof this fleet will have.
|
||||
- **`--acl-storages` semantics are unchanged and the automatic grant does not consult it.** If an
|
||||
operator passes `--acl-storages` deliberately excluding the backup target, the target is still
|
||||
granted by the resolution path. That is the correct precedence — a box cannot function with an
|
||||
unreadable backup target — but it is a place where an override is not absolute, and it is written
|
||||
here rather than left to be discovered.
|
||||
- **`storeGrantRequiredPriv` is a single privilege**, chosen from measurement. If PVE ever changes
|
||||
which privilege gates content listing, the probe would report healthy while the tier is blind. The
|
||||
test asserts the constant's value so a change forces a re-measurement, but nothing detects a change
|
||||
on PVE's side.
|
||||
- **Ten pre-existing `gofmt`-unclean files** remain in the agent repo (unchanged from yesterday's
|
||||
observation); every file touched here is clean.
|
||||
|
||||
@@ -148,6 +148,7 @@
|
||||
| `localapi.DiskOps` / `StorageGate` / `GuestAttacher` / `GuestLister` | internal/localapi/disks.go | `*storage.SudoHostOps`; `storageGateAdapter` (cmd/felhom-agent/main.go); `*GuestBinder`; `*proxmox.Client` | `fakeDiskOps`/`fakeGate`/`fakeGuestAttacher`/`fakeGuestList` internal/localapi/disks_test.go |
|
||||
| `localapi.GuestAPI` / `BackupService` / `BackupStore` / `TokenAuthority` | internal/localapi/server.go | `*proxmox.Client`, `*backup.BackupRunner`, `*backup.Store`, `*TokenStore` | `fakeGuests`/`fakeBackups`/`fakeStore` internal/localapi/server_test.go |
|
||||
| `backup.InFlight` | internal/backup/inflight.go | `TryAcquire(what) (release, busy, ok)` / `Busy()` | THE host-wide "one heavy guest operation at a time" gate — shared by the local-API backup path and the restore-test scheduler (R-85) | A **LINK** guard, not a lock one: the scratch VMID never touches the live guest's vzdump lock, but an offsite restore PULLS multi-GB over the tunnel a backup PUSHES one. Callers **DEFER, never cancel** — a deferred restore-test costs coverage, a cancelled backup costs the backup. A nil gate is ungated (pre-R-85 callers). |
|
||||
| `capability` store-grant probe (`storeGrantStatuses` / `storeGrantVerdict` / `Client.Permissions`) | cmd/felhom-agent/main.go, internal/proxmox/query.go | *"may the agent READ this backup tier?"*, one `capability.Status` per configured tier | R-185. **Never infer permission from an empty content listing** — `{"data":[]}` is what a FORBIDDEN tier and a NEWBORN tier both return, and that ambiguity hid an unreadable host tier on both demo boxes. Ask `/access/permissions` **as the agent's own token** (root always says yes). **The ungranted answer is not empty and not a 403** — it carries the privileges inherited from the box-wide `/` grant, so test for **`Datastore.AllocateSpace`** specifically; path-presence or `Datastore.Audit` reports a blinded storage healthy. Probed set comes from `BackupTiers()`, never a fixed list. Critical except the `local` fallback. Composes AROUND the sudo prober (the `poolReadStatus` precedent); `Status`'s wire shape is untouched so the hub alert is free. Unreachable PVE ⇒ degraded, never ok. |
|
||||
| `backup.RestoreTestState` | internal/backup/restoretest_state.go | `RecordSuccess(target,archive,tier,verified,t)` / `ProvenArchive(target)` / `ProvenRestoreTests(ctx)` / `LastSuccess(target)` / `OldestFirst(targets)` | Per-tier restore-test PROOF state, persisted (atomic tmp+rename) — **which archive** was proven, and when (R-86) | **Credit ONLY on success** — a permanently failing tier must keep sorting first, or it looks freshly proven and stops being retried. Ties break on target id: without it, two tiers proven in the same second rotate by Go's randomised map order. **This one NEEDS persistence unlike R-84** — R-84 had ground truth to consult (the archive is still on the storage); a restore-test destroys its scratch and leaves no artifact. **R-86: the ARCHIVE is the state, the time is metadata** — a time alone cannot answer "have we proven THIS archive", which is the due-check's whole question. A pre-R-86 file (bare RFC3339 per target) keeps its time and yields NO proven archive, so each tier is due once after the upgrade; reading a legacy time as proof of the current archive would invent a guarantee. **R-189: it is also the REPORTABLE half of the restore-test signal.** The in-memory `backup.Store` holds only this process's latest run, and under per-archive due-ness the agent will not re-test a proven archive — so a proof lost to a restart is not repeated for a whole archive generation (observed live: a passing 14.5 GB offsite restore reached no host-report). `ProvenRestoreTests` renders the stored proofs as `hub.RestoreTest` entries and the collector merges them; a record missing the archive or the tier is NOT emitted, because an unproven tier reading as proven is worse than the defect. **Only successes are stored, deliberately:** a success suppresses future work, a failure causes it. |
|
||||
| `hub.ProvenRestoreTestReporter` + `Collector.SetProvenRestoreTests` | internal/hub/collect.go | the DURABLE restore-test source, merged with the in-memory one | R-189. Merge rule: **one entry per tier, newest by `TestedAt` wins** — a fresh failure beats a stored success (the failure is the news, and it lives nowhere else), a stored success beats a stale in-memory entry after a restart, and a tier never appears twice (the hub would read two tests). An unparseable timestamp counts as OLDER, so a malformed entry cannot displace a good one. **The wiring is pinned by an AST test** — the method this replaced (`RestoreTestState.Snapshot`) carried a doc comment naming a host-report gauge and had no caller for weeks. |
|
||||
| `backup.SpecBuilder` / `backup.TierPicker` / `(*BackupRunner).PickSettledRestoreCandidateOn` | internal/backup/schedule.go, runner.go | `func(ctx,archive) RestoreTestSpec`; `func(ctx,target,notAfter) (archive,landed,error)` | The per-run restore-test spec + per-tier **settled** candidate lookup (R-85, widened by R-86) | The spec is built **PER RUN**, never frozen at construction — the pre-R-85 immediately-invoked value made the offsite tier unschedulable AND went stale on any config change. `SourceTier` comes from **the archive**, never the configured target (the v0.100.0 rule). A tier with no archive returns `("", zero, nil)` — **`""` is NOT an error**, or every fresh box looks broken for its first week. **R-86: `notAfter` is the settle cutoff** (zero = no cutoff, which is what keeps `PickRestoreCandidateOn` a one-line call into it), and the picker now skips entries failing `archivePlausiblyComplete` — under per-archive due-ness an incomplete phantom would be picked forever, fail forever, never earn proof, and make the tier due at EVERY evaluation. |
|
||||
|
||||
@@ -0,0 +1,188 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"go/ast"
|
||||
"go/parser"
|
||||
"go/token"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// Scenario H — THE SEAM IS WIRED IN THE PRODUCTION PATH, proven by walking the AST rather than by
|
||||
// grepping for a string.
|
||||
//
|
||||
// WHY THIS TEST EXISTS AND WHY IT IS AN AST WALK. This project's built-but-never-wired count is six,
|
||||
// and links 6 and 7 of the recovery chain were TWO of them: `UnwrapIdentityBundle` sat in the tree
|
||||
// for two months with no caller but a `--selftest`, and the hub's blob-serving endpoints have no
|
||||
// client to this day. The fix must not become the seventh. `strings.Contains` on the file would pass
|
||||
// against a commented-out line, a line inside a test helper, or a line in dead code behind a flag
|
||||
// nobody sets — so this resolves the call graph instead: `Options{EscrowRecovery: …}` must be
|
||||
// constructed inside a function that `runDaemon` reaches, and `runDaemon` must be reached by `main`.
|
||||
|
||||
func parseMain(t *testing.T) (*token.FileSet, *ast.File) {
|
||||
t.Helper()
|
||||
fset := token.NewFileSet()
|
||||
f, err := parser.ParseFile(fset, "main.go", nil, parser.ParseComments)
|
||||
if err != nil {
|
||||
t.Fatalf("parsing main.go: %v", err)
|
||||
}
|
||||
return fset, f
|
||||
}
|
||||
|
||||
// callsWithin returns the set of function names called (directly, by identifier or selector) inside
|
||||
// the named top-level function.
|
||||
func callsWithin(f *ast.File, fnName string) map[string]bool {
|
||||
out := map[string]bool{}
|
||||
for _, d := range f.Decls {
|
||||
fd, ok := d.(*ast.FuncDecl)
|
||||
if !ok || fd.Name == nil || fd.Name.Name != fnName || fd.Body == nil {
|
||||
continue
|
||||
}
|
||||
ast.Inspect(fd.Body, func(n ast.Node) bool {
|
||||
ce, ok := n.(*ast.CallExpr)
|
||||
if !ok {
|
||||
return true
|
||||
}
|
||||
switch fn := ce.Fun.(type) {
|
||||
case *ast.Ident:
|
||||
out[fn.Name] = true
|
||||
case *ast.SelectorExpr:
|
||||
if x, ok := fn.X.(*ast.Ident); ok {
|
||||
out[x.Name+"."+fn.Sel.Name] = true
|
||||
}
|
||||
out[fn.Sel.Name] = true
|
||||
}
|
||||
return true
|
||||
})
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// TestEscrowRecoveryIsWiredIntoTheDaemon asserts the whole chain from func main() to the field.
|
||||
func TestEscrowRecoveryIsWiredIntoTheDaemon(t *testing.T) {
|
||||
_, f := parseMain(t)
|
||||
|
||||
// 1. main() reaches runDaemon.
|
||||
if !callsWithin(f, "main")["runDaemon"] {
|
||||
t.Fatal("func main() does not call runDaemon — the daemon path this test asserts is not the live one")
|
||||
}
|
||||
// 2. runDaemon reaches buildLocalAPIServer.
|
||||
if !callsWithin(f, "runDaemon")["buildLocalAPIServer"] {
|
||||
t.Fatal("runDaemon does not call buildLocalAPIServer — the local API is not built on the daemon path")
|
||||
}
|
||||
|
||||
// 3. Inside buildLocalAPIServer, a localapi.Options composite literal carries EscrowRecovery, and
|
||||
// an escrow.OffsiteKeyRecoverer is constructed there.
|
||||
var optionsHasField, recovererConstructed bool
|
||||
for _, d := range f.Decls {
|
||||
fd, ok := d.(*ast.FuncDecl)
|
||||
if !ok || fd.Name == nil || fd.Name.Name != "buildLocalAPIServer" || fd.Body == nil {
|
||||
continue
|
||||
}
|
||||
ast.Inspect(fd.Body, func(n ast.Node) bool {
|
||||
cl, ok := n.(*ast.CompositeLit)
|
||||
if !ok {
|
||||
return true
|
||||
}
|
||||
sel, ok := cl.Type.(*ast.SelectorExpr)
|
||||
if !ok {
|
||||
return true
|
||||
}
|
||||
pkg, _ := sel.X.(*ast.Ident)
|
||||
if pkg == nil {
|
||||
return true
|
||||
}
|
||||
switch pkg.Name + "." + sel.Sel.Name {
|
||||
case "localapi.Options":
|
||||
for _, el := range cl.Elts {
|
||||
kv, ok := el.(*ast.KeyValueExpr)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
if k, ok := kv.Key.(*ast.Ident); ok && k.Name == "EscrowRecovery" {
|
||||
optionsHasField = true
|
||||
}
|
||||
}
|
||||
case "escrow.OffsiteKeyRecoverer":
|
||||
recovererConstructed = true
|
||||
}
|
||||
return true
|
||||
})
|
||||
}
|
||||
if !recovererConstructed {
|
||||
t.Error("no escrow.OffsiteKeyRecoverer is constructed in buildLocalAPIServer — links 6→8 have no " +
|
||||
"production assembly point (the built-but-never-wired shape, seventh instance)")
|
||||
}
|
||||
if !optionsHasField {
|
||||
t.Error("localapi.Options in buildLocalAPIServer carries no EscrowRecovery field — the recoverer " +
|
||||
"exists and the route would answer 503 forever")
|
||||
}
|
||||
}
|
||||
|
||||
// The hub fetch must be the DAEMON's own hub client, not a freshly constructed one with different
|
||||
// credentials — the self-scoping that makes cross-host retrieval impossible is a property of WHICH
|
||||
// key is used.
|
||||
func TestEscrowRecoveryUsesTheDaemonHubClient(t *testing.T) {
|
||||
fset, f := parseMain(t)
|
||||
var fetchUsesHubClient bool
|
||||
for _, d := range f.Decls {
|
||||
fd, ok := d.(*ast.FuncDecl)
|
||||
if !ok || fd.Name == nil || fd.Name.Name != "buildLocalAPIServer" || fd.Body == nil {
|
||||
continue
|
||||
}
|
||||
ast.Inspect(fd.Body, func(n ast.Node) bool {
|
||||
ce, ok := n.(*ast.CallExpr)
|
||||
if !ok {
|
||||
return true
|
||||
}
|
||||
sel, ok := ce.Fun.(*ast.SelectorExpr)
|
||||
if !ok || sel.Sel.Name != "FetchIdentityEscrow" {
|
||||
return true
|
||||
}
|
||||
if x, ok := sel.X.(*ast.Ident); ok && x.Name == "hubClient" {
|
||||
fetchUsesHubClient = true
|
||||
} else {
|
||||
t.Errorf("FetchIdentityEscrow at %s is called on something other than the injected hub client",
|
||||
fset.Position(ce.Pos()))
|
||||
}
|
||||
return true
|
||||
})
|
||||
}
|
||||
if !fetchUsesHubClient {
|
||||
t.Fatal("the recoverer's fetcher does not call hubClient.FetchIdentityEscrow — either the fetch is " +
|
||||
"not wired, or it uses a client whose credentials are not this host's")
|
||||
}
|
||||
}
|
||||
|
||||
// The route itself must be registered on the local API. A handler with no route is the same defect
|
||||
// one layer down, and it has shipped here before.
|
||||
func TestRecoverRouteIsRegistered(t *testing.T) {
|
||||
fset := token.NewFileSet()
|
||||
f, err := parser.ParseFile(fset, "../../internal/localapi/server.go", nil, 0)
|
||||
if err != nil {
|
||||
t.Fatalf("parsing localapi/server.go: %v", err)
|
||||
}
|
||||
var registered bool
|
||||
ast.Inspect(f, func(n ast.Node) bool {
|
||||
ce, ok := n.(*ast.CallExpr)
|
||||
if !ok || len(ce.Args) < 2 {
|
||||
return true
|
||||
}
|
||||
sel, ok := ce.Fun.(*ast.SelectorExpr)
|
||||
if !ok || sel.Sel.Name != "HandleFunc" {
|
||||
return true
|
||||
}
|
||||
lit, ok := ce.Args[0].(*ast.BasicLit)
|
||||
if !ok {
|
||||
return true
|
||||
}
|
||||
if strings.Contains(lit.Value, "/escrow/recover-offsite-password") {
|
||||
registered = true
|
||||
}
|
||||
return true
|
||||
})
|
||||
if !registered {
|
||||
t.Fatal("POST /escrow/recover-offsite-password is not registered on the local API mux — the handler " +
|
||||
"exists and nothing can reach it")
|
||||
}
|
||||
}
|
||||
+268
-8
@@ -24,6 +24,7 @@ import (
|
||||
"path/filepath"
|
||||
"strconv"
|
||||
"strings"
|
||||
"sync"
|
||||
"syscall"
|
||||
"time"
|
||||
|
||||
@@ -441,15 +442,113 @@ func poolReadStatus(ctx context.Context, px *proxmox.Client) capability.Status {
|
||||
// One exception, so an ordinary configuration is not turned into an alarm: a box with no dedicated
|
||||
// target (`local_backup_target: "local"`, which host-install's own comment calls the DEGRADED
|
||||
// fallback) is not treated as critical for that tier — see storeGrantCritical.
|
||||
func storeGrantStatuses(ctx context.Context, px *proxmox.Client, cfg config.Config) []capability.Status {
|
||||
func storeGrantStatuses(ctx context.Context, px *proxmox.Client, cfg config.Config, repair *storeGrantRepairer) []capability.Status {
|
||||
tiers, _ := cfg.Backup.BackupTiers() // warnings are logged where the tiers are armed
|
||||
out := make([]capability.Status, 0, len(tiers))
|
||||
for _, t := range tiers {
|
||||
out = append(out, storeGrantStatus(ctx, px, t.TargetID, storeGrantCritical(t.TargetID)))
|
||||
out = append(out, storeGrantStatus(ctx, px, t.TargetID, storeGrantCritical(t.TargetID), repair))
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// storeGrantRepairReportWindow is how long after a repair the capability keeps reporting the
|
||||
// transition. It MUST exceed the hub report interval, or the record never reaches the operator.
|
||||
//
|
||||
// FOUND BY THE LIVE RUN, NOT BY THE TESTS (2026-08-04). The first implementation reported degraded
|
||||
// for exactly "one cycle" — the probe call that did the repair. But `probeAll` is invoked
|
||||
// INDEPENDENTLY by the startup/periodic self-check log and by the collector building a host report,
|
||||
// so the repairing call was the LOG's, and the report built three seconds later found the grant
|
||||
// present and reported `ok`. The agent's journal had the record; the hub had nothing; the operator
|
||||
// would have learned nothing. That is precisely the silence R-190 is about, re-created inside its own
|
||||
// mitigation.
|
||||
//
|
||||
// A latch on TIME rather than on call count fixes it: 20 minutes comfortably exceeds the 900 s report
|
||||
// interval, so at least one host-report must carry the transition, and it still clears on its own.
|
||||
const storeGrantRepairReportWindow = 20 * time.Minute
|
||||
|
||||
// storeGrantRepairMinInterval bounds how often a single tier's grant may be re-granted (Scenario F).
|
||||
//
|
||||
// A storage can be unreadable for reasons an ACL cannot fix — the storage is gone, PVE is wedged,
|
||||
// the wrapper is missing. Without a bound the probe would re-grant on every report cycle forever: a
|
||||
// repair loop is a new defect wearing a fix's clothes. One attempt per tier per hour is frequent
|
||||
// enough that a real loss is repaired within one backup window, and rare enough that a permanent
|
||||
// fault produces attempts you can count on one hand per day.
|
||||
const storeGrantRepairMinInterval = time.Hour
|
||||
|
||||
// storeGrantRepairer bounds and records the self-repair. It is deliberately in-memory: an agent
|
||||
// restart re-arms the repair, which is correct — a restart is exactly when a box should re-check
|
||||
// everything it depends on.
|
||||
type storeGrantRepairer struct {
|
||||
run func(ctx context.Context, name string, args ...string) ([]byte, []byte, error)
|
||||
log *slog.Logger
|
||||
mu sync.Mutex
|
||||
last map[string]time.Time // target id → last ATTEMPT (success or failure)
|
||||
repaired map[string]time.Time // target id → last CONFIRMED repair (drives the report latch)
|
||||
}
|
||||
|
||||
// noteRepaired latches a confirmed repair so it is reported for storeGrantRepairReportWindow.
|
||||
func (r *storeGrantRepairer) noteRepaired(target string, now time.Time) {
|
||||
if r == nil {
|
||||
return
|
||||
}
|
||||
r.mu.Lock()
|
||||
defer r.mu.Unlock()
|
||||
if r.repaired == nil {
|
||||
r.repaired = map[string]time.Time{}
|
||||
}
|
||||
r.repaired[target] = now
|
||||
}
|
||||
|
||||
// recentlyRepaired reports whether a confirmed repair is still inside its report window — the latch
|
||||
// that guarantees a host-report carries the transition even though the probe that repaired may have
|
||||
// been a log-only one.
|
||||
func (r *storeGrantRepairer) recentlyRepaired(target string, now time.Time) bool {
|
||||
if r == nil {
|
||||
return false
|
||||
}
|
||||
r.mu.Lock()
|
||||
defer r.mu.Unlock()
|
||||
t, ok := r.repaired[target]
|
||||
return ok && now.Sub(t) < storeGrantRepairReportWindow
|
||||
}
|
||||
|
||||
// mayAttempt reports whether a repair may run now for this target, and records the attempt if so.
|
||||
func (r *storeGrantRepairer) mayAttempt(target string, now time.Time) bool {
|
||||
if r == nil || r.run == nil {
|
||||
return false
|
||||
}
|
||||
r.mu.Lock()
|
||||
defer r.mu.Unlock()
|
||||
if r.last == nil {
|
||||
r.last = map[string]time.Time{}
|
||||
}
|
||||
if t, ok := r.last[target]; ok && now.Sub(t) < storeGrantRepairMinInterval {
|
||||
return false
|
||||
}
|
||||
r.last[target] = now
|
||||
return true
|
||||
}
|
||||
|
||||
// repair runs the EXISTING root wrapper's `grant` verb for this storage. It adds no privileged
|
||||
// surface: `felhom-backup-target-apply grant *` is already in the sudoers allowlist for any storage
|
||||
// id (configs/felhom-agent.sudoers), and the verb already grants BOTH the user and the token — a
|
||||
// privsep token's rights are the intersection, so granting one of the two grants nothing usable.
|
||||
//
|
||||
// This is the pbsdr shape (internal/pbsdr/manager.go, the R-22 self-grant): on a refusal, run the
|
||||
// root wrapper and RE-READ ONCE rather than dead-locking. Its restraint is copied too — one attempt,
|
||||
// one confirmation, and anything still wrong stays loudly wrong.
|
||||
func (r *storeGrantRepairer) repair(ctx context.Context, target string) error {
|
||||
rctx, cancel := context.WithTimeout(ctx, 30*time.Second)
|
||||
defer cancel()
|
||||
_, errOut, err := r.run(rctx, localapi.BackupTargetWrapperPath, "grant", target)
|
||||
if err != nil {
|
||||
r.log.Error("store-grant: SELF-REPAIR FAILED — the tier stays unreadable",
|
||||
"target", target, "err", err, "stderr", strings.TrimSpace(string(errOut)))
|
||||
return err
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// storeGrantRequiredPriv is the privilege whose ABSENCE was measured to blind the content listing.
|
||||
//
|
||||
// Measured on demo-felhom 2026-08-03: the two storages that list through the token hold
|
||||
@@ -468,7 +567,7 @@ func storeGrantCritical(targetID string) bool { return targetID != "local" }
|
||||
|
||||
// storeGrantStatus is one tier's grant probe. It NEVER reports ok when it could not ask: a
|
||||
// self-check that fails open is worse than none, because it converts "I do not know" into "fine".
|
||||
func storeGrantStatus(ctx context.Context, px *proxmox.Client, targetID string, critical bool) capability.Status {
|
||||
func storeGrantStatus(ctx context.Context, px *proxmox.Client, targetID string, critical bool, repair *storeGrantRepairer) capability.Status {
|
||||
s := capability.Status{
|
||||
Name: "pve:store-grant:" + targetID,
|
||||
Feature: "backup tier " + targetID + " readable by the agent (archive listing, restore-test candidacy)",
|
||||
@@ -486,7 +585,117 @@ func storeGrantStatus(ctx context.Context, px *proxmox.Client, targetID string,
|
||||
pctx, cancel := context.WithTimeout(ctx, 10*time.Second)
|
||||
defer cancel()
|
||||
privs, err := px.Permissions(pctx, "/storage/"+targetID)
|
||||
return storeGrantVerdict(targetID, critical, privs, err)
|
||||
s = storeGrantVerdict(targetID, critical, privs, err)
|
||||
if err != nil {
|
||||
return s
|
||||
}
|
||||
if s.Status != capability.StatusDegraded {
|
||||
// Healthy — but if this tier was repaired moments ago, keep REPORTING the transition until a
|
||||
// host-report has certainly carried it. Without this latch the repairing probe may be a
|
||||
// log-only one and the hub never learns anything happened (measured live, see the window's
|
||||
// comment).
|
||||
return storeGrantHealthyVerdict(targetID, critical, s, repair.recentlyRepaired(targetID, time.Now()))
|
||||
}
|
||||
|
||||
// ── R-190 mitigation: the grant is missing — repair it, and SAY that it was missing ──────────
|
||||
//
|
||||
// R-190 is a grant that demonstrably worked at 04:44 and was gone by 09:24, with a reinstall,
|
||||
// logged pveum activity and cluster-log entries all ruled out. The cause is still open; the
|
||||
// resilience does not have to wait for it. Everything needed already exists — the root wrapper,
|
||||
// its sudoers vector for any storage id, and the exact command — and until now the `grant` verb
|
||||
// had only ever been called at CREATION. That is the "built but never wired" shape, in a verb
|
||||
// rather than a seam.
|
||||
if !repair.mayAttempt(targetID, time.Now()) {
|
||||
// Bounded (Scenario F): an earlier attempt did not hold and it is too soon to try again. Stay
|
||||
// degraded and say why — a quiet "we already tried" is how a permanent fault becomes silence.
|
||||
s.Reason = "the agent token lacks " + storeGrantRequiredPriv + " on /storage/" + targetID +
|
||||
" and a self-repair was attempted within the last " + storeGrantRepairMinInterval.String() +
|
||||
" without holding — NOT retrying yet; this needs a human"
|
||||
return s
|
||||
}
|
||||
if rerr := repair.repair(ctx, targetID); rerr != nil {
|
||||
s.Reason = "the agent token lacks " + storeGrantRequiredPriv + " on /storage/" + targetID +
|
||||
" and the self-repair FAILED (" + rerr.Error() + ") — this tier's archives are INVISIBLE to the agent"
|
||||
return s // Scenario E: a failed repair must never mask the degraded state.
|
||||
}
|
||||
// Re-read ONCE to confirm, exactly as pbsdr does — the wrapper reporting success is a claim about
|
||||
// its own write; the grant being readable is a different claim, and it is the one that matters.
|
||||
cctx, ccancel := context.WithTimeout(ctx, 10*time.Second)
|
||||
defer ccancel()
|
||||
privs2, err2 := px.Permissions(cctx, "/storage/"+targetID)
|
||||
if err2 != nil || privs2[storeGrantRequiredPriv] != 1 {
|
||||
s.Reason = "the agent token lacks " + storeGrantRequiredPriv + " on /storage/" + targetID +
|
||||
" and the self-repair did not take (re-read says it is still missing) — this needs a human"
|
||||
return s
|
||||
}
|
||||
|
||||
// REPAIRED — and reported as DEGRADED for exactly this one cycle, deliberately.
|
||||
//
|
||||
// The tier works again, so "ok" would be true of this instant and would throw away the only
|
||||
// evidence that anything happened. R-190's own words: the probe sees the STATE, nothing sees the
|
||||
// TRANSITION. A silent self-repair makes a recurring loss undetectable forever, which is strictly
|
||||
// worse than the fault it fixes.
|
||||
//
|
||||
// §8.5 asked whether the hub's existing degraded↔ok edge suffices before building anything new.
|
||||
// It does — as a CHANNEL — but only if the agent deliberately reports one degraded cycle: the hub
|
||||
// alerts and e-mails on the ok→degraded edge and logs the degraded→ok recovery, so one loss
|
||||
// produces exactly one alert pair and the operator learns of it. NOTHING NEW WAS BUILT: no wire
|
||||
// change, no hub change, no new event type. The `Feature` text carries the explanation because
|
||||
// that is the field the hub puts in the operator's e-mail (the Reason does not travel).
|
||||
repair.noteRepaired(targetID, time.Now())
|
||||
s = storeGrantRepairedVerdict(targetID, critical)
|
||||
repairLogger(repair).Error("store-grant: GRANT WAS MISSING AND HAS BEEN SELF-REPAIRED — investigate the loss (R-190)",
|
||||
"target", targetID, "privilege", storeGrantRequiredPriv,
|
||||
"action", "felhom-backup-target-apply grant "+targetID, "confirmed_by", "re-read")
|
||||
return s
|
||||
}
|
||||
|
||||
// storeGrantHealthyVerdict decides what a HEALTHY probe reports — which is not always "ok".
|
||||
//
|
||||
// Split out so the tests exercise this decision rather than a copy of it. An earlier version of this
|
||||
// guard lived inline and its red-proof PASSED, because the test asserted the latch helper instead of
|
||||
// the path that consumes it — the same hollow shape this file has now caught twice.
|
||||
//
|
||||
// If the tier was repaired inside the report window, the transition is reported even though the grant
|
||||
// is present: the probe that repaired may have been a log-only one, and without this the host-report
|
||||
// carries `ok` and the operator never learns the permission vanished (measured live 2026-08-04).
|
||||
func storeGrantHealthyVerdict(targetID string, critical bool, healthy capability.Status, repairedRecently bool) capability.Status {
|
||||
if repairedRecently {
|
||||
return storeGrantRepairedVerdict(targetID, critical)
|
||||
}
|
||||
return healthy
|
||||
}
|
||||
|
||||
// storeGrantRepairedVerdict is the post-repair verdict — the RECORD half of R-190, split out so the
|
||||
// tests exercise the real thing rather than a copy of it (yesterday's hollow-test lesson).
|
||||
//
|
||||
// It reports DEGRADED although the tier now works, and that is the whole point: "ok" would be true of
|
||||
// this instant and would throw away the only evidence that a permission vanished. The hub raises its
|
||||
// ok→degraded edge (an operator e-mail) and logs the degraded→ok recovery on the next cycle, so one
|
||||
// loss produces exactly one alert pair. Nothing new was built for this — no wire change, no hub
|
||||
// change, no new event type.
|
||||
//
|
||||
// The explanation lives in FEATURE because that is the field the hub interpolates into the operator's
|
||||
// e-mail (`monitor/host_capability.go` emitTransition builds its message from the capability names
|
||||
// and features; Reason does not travel). Putting it in Reason alone would be a record nobody reads.
|
||||
func storeGrantRepairedVerdict(targetID string, critical bool) capability.Status {
|
||||
return capability.Status{
|
||||
Name: "pve:store-grant:" + targetID,
|
||||
Critical: critical,
|
||||
Status: capability.StatusDegraded,
|
||||
Feature: "backup tier " + targetID + ": the agent's storage grant was MISSING and has been " +
|
||||
"AUTOMATICALLY RESTORED — the tier works now, but a permission that vanished on its own needs investigating (R-190)",
|
||||
Reason: "grant absent at probe time; `felhom-backup-target-apply grant " + targetID +
|
||||
"` re-applied it and a re-read confirms " + storeGrantRequiredPriv + " is present again",
|
||||
}
|
||||
}
|
||||
|
||||
// repairLogger returns the repairer's logger, or the default — the record must survive a nil.
|
||||
func repairLogger(r *storeGrantRepairer) *slog.Logger {
|
||||
if r != nil && r.log != nil {
|
||||
return r.log
|
||||
}
|
||||
return slog.Default()
|
||||
}
|
||||
|
||||
// storeGrantVerdict is the DECISION, split out from the API call so the tests exercise the real
|
||||
@@ -598,9 +807,17 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
|
||||
// reaper fail-safes (locks stay uncleared) — visible on the hub report, no operator page.
|
||||
// R-185: the store-grant probes compose around the sudo prober the same way the pool read does
|
||||
// (an API read does not belong inside the sudo-policy probe — the v0.62.0 A1 precedent).
|
||||
// R-190: the store-grant probe also REPAIRS a missing grant, through the root wrapper that
|
||||
// already exists and is already sudoers-permitted for any storage id — and reports the loss.
|
||||
// The runner is the DIRECT one for the same reason the sudo prober uses it: the wrapper is
|
||||
// invoked through the privileged path, which prepends sudo itself.
|
||||
grantRepairer := &storeGrantRepairer{
|
||||
run: (&proxmox.ExecRunner{Mode: proxmox.RunnerMode(cfg.Privileged.Mode)}).Run,
|
||||
log: logger,
|
||||
}
|
||||
probeAll := func(ctx context.Context) []capability.Status {
|
||||
out := append(capProber.Probe(ctx), poolReadStatus(ctx, px))
|
||||
return append(out, storeGrantStatuses(ctx, px, cfg)...)
|
||||
return append(out, storeGrantStatuses(ctx, px, cfg, grantRepairer)...)
|
||||
}
|
||||
// (The startup self-check log runs AFTER the pbsdr manager is wired below, so its snapshot
|
||||
// already carries the gated view — v0.86.0.)
|
||||
@@ -882,7 +1099,7 @@ func runDaemon(cfg config.Config, logger *slog.Logger, logRing *applog.Ring) int
|
||||
return false
|
||||
},
|
||||
}
|
||||
localSrv := buildLocalAPIServer(cfg, px, backupStore, heavyOps, observer, driveKnown, hostOps, gate, collector, intentRec, guestBindStore, formatJobStore, logRing, escrowCeremonyCfg, logger, &localTokens)
|
||||
localSrv := buildLocalAPIServer(cfg, px, backupStore, heavyOps, observer, driveKnown, hostOps, gate, collector, client, intentRec, guestBindStore, formatJobStore, logRing, escrowCeremonyCfg, logger, &localTokens)
|
||||
if localTokens != nil {
|
||||
defer localTokens.Close()
|
||||
}
|
||||
@@ -1460,7 +1677,7 @@ func buildRestoreTestScheduler(cfg config.Config, px *proxmox.Client, engine *re
|
||||
// leaf (stable fingerprint). Any failure DISABLES the server (returns nil) WITHOUT crashing the
|
||||
// daemon — the host still reports/reconciles; only the controller channel is unavailable until
|
||||
// fixed. The opened token store is returned via outTokens so the caller can Close it.
|
||||
func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.Store, inFlight *backup.InFlight, observer *storage.Observer, driveTargets storage.KnownTargets, hostOps *storage.SudoHostOps, gate *reconcile.Gate, collector *hub.Collector, intent localapi.IntentRecorder, guestBinds *localapi.GuestBindStore, formatJobs *localapi.FormatJobStore, logRing *applog.Ring, escrowCeremony *localapi.EscrowCeremonyConfig, logger *slog.Logger, outTokens **localapi.TokenStore) *localapi.Server {
|
||||
func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.Store, inFlight *backup.InFlight, observer *storage.Observer, driveTargets storage.KnownTargets, hostOps *storage.SudoHostOps, gate *reconcile.Gate, collector *hub.Collector, hubClient *hub.Client, intent localapi.IntentRecorder, guestBinds *localapi.GuestBindStore, formatJobs *localapi.FormatJobStore, logRing *applog.Ring, escrowCeremony *localapi.EscrowCeremonyConfig, logger *slog.Logger, outTokens **localapi.TokenStore) *localapi.Server {
|
||||
if !cfg.LocalAPI.Enabled() {
|
||||
return nil
|
||||
}
|
||||
@@ -1532,7 +1749,29 @@ func buildLocalAPIServer(cfg config.Config, px *proxmox.Client, store *backup.St
|
||||
gaMode = proxmox.RunnerSudo
|
||||
}
|
||||
guestBinder := localapi.NewGuestBinder(&proxmox.ExecRunner{Mode: gaMode, SudoPath: cfg.Privileged.SudoPath}, logger)
|
||||
// R-199 (v0.125.0) — chain links 6->8, assembled here and ONLY here. The fetcher is this daemon's
|
||||
// own hub client (per-host key, self-scoped server-side), so the recoverer can never read another
|
||||
// host's blob even if asked to. `client` is the same one the report loop uses; a nil hub config
|
||||
// cannot reach this line (the daemon exits above), so the seam is always live in production —
|
||||
// which is the point: links 6 and 7 spent months existing without a caller.
|
||||
escrowRecoverer := escrow.OffsiteKeyRecoverer{
|
||||
Fetch: func(ctx context.Context) ([]byte, bool, error) {
|
||||
resp, ferr := hubClient.FetchIdentityEscrow(ctx)
|
||||
if ferr != nil {
|
||||
return nil, false, ferr
|
||||
}
|
||||
if !resp.Present || resp.IdentityEscrowB64 == "" {
|
||||
return nil, false, nil
|
||||
}
|
||||
blob, derr := base64.StdEncoding.DecodeString(resp.IdentityEscrowB64)
|
||||
if derr != nil {
|
||||
return nil, false, fmt.Errorf("hub served a malformed escrow blob (not base64)")
|
||||
}
|
||||
return blob, true, nil
|
||||
},
|
||||
}
|
||||
srv, err := localapi.NewServer(localapi.Options{
|
||||
EscrowRecovery: escrowRecoverer,
|
||||
ListenAddr: cfg.LocalAPI.ListenAddr,
|
||||
Cert: cert,
|
||||
AgentVersion: version, // v0.82.0: the X-Felhom-Agent-Version capability channel
|
||||
@@ -2654,7 +2893,28 @@ func runSelftestIdentityConsume(ctx context.Context, cfg config.Config, logger *
|
||||
fmt.Fprintln(os.Stderr, " [FAIL] writing recovered bundle:", err)
|
||||
return 1
|
||||
}
|
||||
fmt.Printf(" [OK] identity recovered (tunnel_token + pbs_token) → %s (0600) — never printed\n", keyDest)
|
||||
// R-199 / §8.6: this line used to read "(tunnel_token + pbs_token)" — an enumeration that was
|
||||
// accurate when it was written (pre-fork-4) and became a MISSTATEMENT the moment v0.77.0 sealed the
|
||||
// offsite repository password into the same bundle. Anyone reading the old output would conclude the
|
||||
// repository password was not there, and that is part of how the chain's extraction link came to be
|
||||
// described as missing for a month. Name what was recovered from THIS bundle, and name what is
|
||||
// absent, rather than reciting a fixed list.
|
||||
recovered := []string{"tunnel_token", "pbs_token"}
|
||||
var absent []string
|
||||
if bundle.WGPrivateKey != "" {
|
||||
recovered = append(recovered, "wg_private_key")
|
||||
} else {
|
||||
absent = append(absent, "wg_private_key")
|
||||
}
|
||||
if bundle.ResticRepoPassword != "" {
|
||||
recovered = append(recovered, "restic_repo_password")
|
||||
} else {
|
||||
absent = append(absent, "restic_repo_password (pre-fork-4 blob — the field did not exist when this was sealed)")
|
||||
}
|
||||
fmt.Printf(" [OK] identity recovered (%s) → %s (0600) — values never printed\n", strings.Join(recovered, " + "), keyDest)
|
||||
if len(absent) > 0 {
|
||||
fmt.Printf(" [NOTE] fields ABSENT from this bundle: %s\n", strings.Join(absent, "; "))
|
||||
}
|
||||
|
||||
// S5 DR: install the recovered WG private key so the tunnel re-establishes with the SAME
|
||||
// identity/pubkey (→ the same hub /32), no fresh keygen. Create-only (refuses to overwrite a
|
||||
|
||||
@@ -2,9 +2,13 @@ package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"go/ast"
|
||||
"io"
|
||||
"log/slog"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/capability"
|
||||
)
|
||||
@@ -132,7 +136,7 @@ func TestStoreGrant_TheFallbackTargetIsNotCritical(t *testing.T) {
|
||||
// A probe that cannot ask must never answer "ok" — unknown reported as healthy is worse than no
|
||||
// probe, because it looks like coverage.
|
||||
func TestStoreGrant_UnreachablePVEIsDegradedNotOK(t *testing.T) {
|
||||
s := storeGrantStatus(context.Background(), nil, "felhom-backup", true)
|
||||
s := storeGrantStatus(context.Background(), nil, "felhom-backup", true, nil)
|
||||
if s.Status != capability.StatusDegraded {
|
||||
t.Fatalf("an unaskable probe must be DEGRADED, never ok; got %q", s.Status)
|
||||
}
|
||||
@@ -164,3 +168,250 @@ func TestMainWiresTheStoreGrantProbe(t *testing.T) {
|
||||
"which is precisely the silence R-185 is about")
|
||||
}
|
||||
}
|
||||
|
||||
// ── R-190 — the grant repairs itself, and the repair is VISIBLE ──────────────────────────────
|
||||
//
|
||||
// R-190 is a storage grant that demonstrably worked at 04:44 on 2026-08-03 and was gone by 09:24,
|
||||
// with a host reinstall, logged `pveum` activity and cluster-log entries all ruled out. The cause is
|
||||
// open; the resilience is not conditional on it.
|
||||
//
|
||||
// The half that matters is the RECORD. R-190's own words: the probe sees the state, nothing sees the
|
||||
// transition. A self-repair that leaves only "ok" behind destroys the only evidence a loss happened,
|
||||
// so a recurring loss becomes undetectable forever — strictly worse than the fault it fixes.
|
||||
|
||||
// fakeRepairRunner records wrapper invocations and can be made to fail.
|
||||
type fakeRepairRunner struct {
|
||||
calls [][]string
|
||||
fail bool
|
||||
}
|
||||
|
||||
func (f *fakeRepairRunner) Run(_ context.Context, name string, args ...string) ([]byte, []byte, error) {
|
||||
f.calls = append(f.calls, append([]string{name}, args...))
|
||||
if f.fail {
|
||||
return nil, []byte("pveum: refused"), errors.New("exit status 2")
|
||||
}
|
||||
return nil, nil, nil
|
||||
}
|
||||
|
||||
func newRepairer(f *fakeRepairRunner) *storeGrantRepairer {
|
||||
return &storeGrantRepairer{run: f.Run, log: slog.New(slog.NewTextHandler(io.Discard, nil))}
|
||||
}
|
||||
|
||||
// ── SCENARIO F — the repair is BOUNDED ───────────────────────────────────────────────────────
|
||||
//
|
||||
// COMPANION RED-PROOF (observed 2026-08-04): make mayAttempt always return true (drop the
|
||||
// storeGrantRepairMinInterval check) →
|
||||
//
|
||||
// --- FAIL: TestGrantRepair_IsBounded
|
||||
// storegrant_test.go: a repair must not run on every cycle; 5 cycles produced 5 attempt(s)
|
||||
//
|
||||
// which is a re-grant every report cycle, forever, against a fault an ACL cannot fix. Restored.
|
||||
func TestGrantRepair_IsBounded(t *testing.T) {
|
||||
f := &fakeRepairRunner{}
|
||||
r := newRepairer(f)
|
||||
// Jittered, so the series never lands exactly on the interval boundary — a perfectly regular
|
||||
// series is how a threshold test passes its own mutation, which has happened here before.
|
||||
base := time.Date(2026, 8, 4, 9, 17, 43, 0, time.UTC)
|
||||
offsets := []time.Duration{0, 13*time.Minute + 7*time.Second, 27*time.Minute + 51*time.Second,
|
||||
41*time.Minute + 19*time.Second, 55*time.Minute + 3*time.Second}
|
||||
attempts := 0
|
||||
for _, off := range offsets {
|
||||
if r.mayAttempt("felhom-backup", base.Add(off)) {
|
||||
attempts++
|
||||
}
|
||||
}
|
||||
if attempts != 1 {
|
||||
t.Fatalf("a repair must not run on every cycle; %d cycles produced %d attempt(s) within %s",
|
||||
len(offsets), attempts, storeGrantRepairMinInterval)
|
||||
}
|
||||
// ...and once the interval has genuinely passed, it may try again — a bound is not a ban.
|
||||
if !r.mayAttempt("felhom-backup", base.Add(storeGrantRepairMinInterval+2*time.Minute+11*time.Second)) {
|
||||
t.Fatal("after the interval a repair must be allowed again — otherwise one failure disables the repair forever")
|
||||
}
|
||||
// A DIFFERENT tier is not throttled by this one's attempt.
|
||||
if !r.mayAttempt("felhom-pbs", base.Add(time.Minute)) {
|
||||
t.Fatal("the bound must be per tier — one tier's attempt must not suppress another's")
|
||||
}
|
||||
}
|
||||
|
||||
// A nil repairer (or one with no runner) never attempts, and never panics.
|
||||
func TestGrantRepair_NilIsSafe(t *testing.T) {
|
||||
var r *storeGrantRepairer
|
||||
if r.mayAttempt("felhom-backup", time.Now()) {
|
||||
t.Fatal("a nil repairer must never claim an attempt")
|
||||
}
|
||||
if (&storeGrantRepairer{}).mayAttempt("felhom-backup", time.Now()) {
|
||||
t.Fatal("a repairer with no runner must never claim an attempt")
|
||||
}
|
||||
}
|
||||
|
||||
// The repair calls the EXISTING wrapper verb, with the storage id — no new privileged surface.
|
||||
func TestGrantRepair_CallsTheExistingWrapperVerb(t *testing.T) {
|
||||
f := &fakeRepairRunner{}
|
||||
r := newRepairer(f)
|
||||
if err := r.repair(context.Background(), "felhom-backup"); err != nil {
|
||||
t.Fatalf("repair should succeed with a healthy runner: %v", err)
|
||||
}
|
||||
if len(f.calls) != 1 {
|
||||
t.Fatalf("exactly one wrapper invocation expected; got %d", len(f.calls))
|
||||
}
|
||||
got := f.calls[0]
|
||||
want := []string{"/usr/local/sbin/felhom-backup-target-apply", "grant", "felhom-backup"}
|
||||
if len(got) != len(want) {
|
||||
t.Fatalf("wrapper argv = %v, want %v", got, want)
|
||||
}
|
||||
for i := range want {
|
||||
if got[i] != want[i] {
|
||||
t.Fatalf("wrapper argv = %v, want %v — the sudoers vector is `grant *`; anything else is a policy change", got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A repair that FAILS must surface the failure, not swallow it (Scenario E's precondition).
|
||||
func TestGrantRepair_FailureIsReturned(t *testing.T) {
|
||||
f := &fakeRepairRunner{fail: true}
|
||||
if err := newRepairer(f).repair(context.Background(), "felhom-backup"); err == nil {
|
||||
t.Fatal("a failed wrapper run must return its error — a repair that cannot run must never read as done")
|
||||
}
|
||||
}
|
||||
|
||||
// ── SCENARIO D (the half that matters) — the REPAIR MUST BE VISIBLE ──────────────────────────
|
||||
//
|
||||
// A repair that leaves only "ok" behind is worse than the fault: the tier works, and the fact that a
|
||||
// permission vanished is gone with it. R-190 exists because nothing saw the transition.
|
||||
//
|
||||
// The channel is the hub's EXISTING ok→degraded→ok edge (§8.5) — nothing new was built. That only
|
||||
// works if the agent deliberately reports ONE degraded cycle after repairing, and if the explanation
|
||||
// rides the field the hub actually puts in the operator's e-mail. The hub's message is built from the
|
||||
// capability NAME and FEATURE (`internal/monitor/host_capability.go` emitTransition) — **not** from
|
||||
// Reason — so the Feature must carry it.
|
||||
//
|
||||
// COMPANION RED-PROOF (observed 2026-08-04): after a successful repair, report ok instead —
|
||||
//
|
||||
// s.Status = capability.StatusOK; s.Feature unchanged
|
||||
//
|
||||
// → --- FAIL: TestGrantRepair_ARepairedGrantIsReportedAsATransition
|
||||
//
|
||||
// storegrant_test.go: a self-repair must still report DEGRADED for one cycle so the hub raises
|
||||
// its edge; got "ok" — the loss would be invisible
|
||||
//
|
||||
// i.e. exactly the silence R-190 is about. Restored.
|
||||
func TestGrantRepair_ARepairedGrantIsReportedAsATransition(t *testing.T) {
|
||||
// THE PRODUCTION verdict, not a copy of it. An earlier draft of this test built the Status
|
||||
// itself and asserted its own construction — it would have passed while production reported ok,
|
||||
// which is precisely the silence being guarded against.
|
||||
if pre := probeWith(permUngranted, "felhom-backup", true); pre.Status != capability.StatusDegraded {
|
||||
t.Fatalf("precondition: a missing grant is degraded; got %q", pre.Status)
|
||||
}
|
||||
s := storeGrantRepairedVerdict("felhom-backup", true)
|
||||
|
||||
if s.Status != capability.StatusDegraded {
|
||||
t.Fatalf("a self-repair must still report DEGRADED for one cycle so the hub raises its edge; "+
|
||||
"got %q — the loss would be invisible", s.Status)
|
||||
}
|
||||
// The hub e-mails the FEATURE text. If the explanation is not there, the operator is told a
|
||||
// capability was degraded and never learns it repaired itself or that anything vanished.
|
||||
for _, want := range []string{"MISSING", "RESTORED", "felhom-backup", "R-190"} {
|
||||
if !strings.Contains(s.Feature, want) {
|
||||
t.Fatalf("the Feature text is what the hub puts in the operator's e-mail; it must contain %q. Got: %s", want, s.Feature)
|
||||
}
|
||||
}
|
||||
if !s.Critical {
|
||||
t.Fatal("the transition must be CRITICAL or the hub does not alert on it at all")
|
||||
}
|
||||
}
|
||||
|
||||
// ── SCENARIO H — the seam ────────────────────────────────────────────────────────────────────
|
||||
//
|
||||
// The wrapper's `grant` verb is itself a "built but never wired" example: it exists, is
|
||||
// sudoers-permitted for any id, and had only ever been called at storage CREATION. The repair must
|
||||
// not become the seventh instance. AST, not grep — a commented-out call still contains the string.
|
||||
func TestMainWiresTheGrantRepair(t *testing.T) {
|
||||
f := parseMainForWiring(t)
|
||||
|
||||
var built, passed bool
|
||||
ast.Inspect(f, func(n ast.Node) bool {
|
||||
switch node := n.(type) {
|
||||
case *ast.CompositeLit:
|
||||
if id, ok := node.Type.(*ast.Ident); ok && id.Name == "storeGrantRepairer" {
|
||||
built = true
|
||||
}
|
||||
case *ast.CallExpr:
|
||||
if id, ok := node.Fun.(*ast.Ident); ok && id.Name == "storeGrantStatuses" && len(node.Args) == 4 {
|
||||
if a, ok := node.Args[3].(*ast.Ident); ok && a.Name == "grantRepairer" {
|
||||
passed = true
|
||||
}
|
||||
}
|
||||
}
|
||||
return true
|
||||
})
|
||||
if !built {
|
||||
t.Error("main.go never constructs a storeGrantRepairer — nothing would ever repair a lost grant")
|
||||
}
|
||||
if !passed {
|
||||
t.Error("storeGrantStatuses is not passed the repairer — the probe would detect the loss and " +
|
||||
"leave it, which is v0.123.0's behaviour and not R-190's mitigation")
|
||||
}
|
||||
}
|
||||
|
||||
// The transition must survive a probe that is NOT the one feeding the hub.
|
||||
//
|
||||
// MEASURED LIVE 2026-08-04, and this test exists because the first implementation failed it in
|
||||
// production while every unit test passed: `probeAll` is called independently by the self-check LOG
|
||||
// and by the collector building a host-report. The repairing call was the log's; the report three
|
||||
// seconds later found the grant present and reported `ok`. The agent's journal had the record and the
|
||||
// hub had nothing — the exact silence R-190 is about, re-created inside its own mitigation.
|
||||
//
|
||||
// COMPANION RED-PROOF (observed): delete the `recentlyRepaired` branch from the healthy path →
|
||||
//
|
||||
// --- FAIL: TestGrantRepair_TransitionSurvivesALaterProbe
|
||||
// storegrant_test.go: a probe AFTER the repair must still report the transition; got "ok" —
|
||||
// the host-report would carry ok and the operator would never learn the grant vanished
|
||||
//
|
||||
// Restored.
|
||||
func TestGrantRepair_TransitionSurvivesALaterProbe(t *testing.T) {
|
||||
r := newRepairer(&fakeRepairRunner{})
|
||||
// Jittered, never landing on the window boundary.
|
||||
repairedAt := time.Date(2026, 8, 4, 9, 39, 34, 0, time.UTC)
|
||||
r.noteRepaired("felhom-backup", repairedAt)
|
||||
|
||||
// The DECISION a later probe makes — the production function, not the helper it calls. An
|
||||
// earlier draft asserted `recentlyRepaired` directly and its red-proof PASSED, because removing
|
||||
// the latch's USE left the helper untouched.
|
||||
healthy := probeWith(permGranted, "felhom-backup", true)
|
||||
if healthy.Status != capability.StatusOK {
|
||||
t.Fatalf("precondition: a granted tier is ok; got %q", healthy.Status)
|
||||
}
|
||||
got := storeGrantHealthyVerdict("felhom-backup", true,
|
||||
healthy, r.recentlyRepaired("felhom-backup", repairedAt.Add(3*time.Second)))
|
||||
if got.Status != capability.StatusDegraded {
|
||||
t.Fatalf("a probe AFTER the repair must still report the transition; got %q — the host-report "+
|
||||
"would carry ok and the operator would never learn the grant vanished", got.Status)
|
||||
}
|
||||
if !strings.Contains(got.Feature, "RESTORED") {
|
||||
t.Fatalf("the later probe must carry the explanation into the hub's e-mail; got: %s", got.Feature)
|
||||
}
|
||||
// Outside the window it reports plain ok again.
|
||||
late := storeGrantHealthyVerdict("felhom-backup", true,
|
||||
healthy, r.recentlyRepaired("felhom-backup", repairedAt.Add(storeGrantRepairReportWindow+time.Minute)))
|
||||
if late.Status != capability.StatusOK {
|
||||
t.Fatalf("outside the window a healthy tier reports ok; got %q — a permanent degraded state "+
|
||||
"would be its own false alarm", late.Status)
|
||||
}
|
||||
if !r.recentlyRepaired("felhom-backup", repairedAt.Add(14*time.Minute+37*time.Second)) {
|
||||
t.Fatal("the latch must outlast the 900s hub report interval, or the record never reaches the hub")
|
||||
}
|
||||
// ...and it clears on its own rather than latching a box degraded forever.
|
||||
if r.recentlyRepaired("felhom-backup", repairedAt.Add(storeGrantRepairReportWindow+time.Minute+7*time.Second)) {
|
||||
t.Fatal("the latch must clear — a permanent degraded state would be its own false alarm")
|
||||
}
|
||||
// It is per tier.
|
||||
if r.recentlyRepaired("felhom-pbs", repairedAt.Add(time.Second)) {
|
||||
t.Fatal("one tier's repair must not latch another tier's status")
|
||||
}
|
||||
// The window MUST exceed the report interval — the property, asserted rather than assumed.
|
||||
if storeGrantRepairReportWindow <= 15*time.Minute {
|
||||
t.Fatalf("the report window (%s) must exceed the 900s hub report interval, or a transition can "+
|
||||
"be missed entirely", storeGrantRepairReportWindow)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,74 @@
|
||||
package escrow
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
)
|
||||
|
||||
// R-199 links 6→8 — fetch this host's own sealed identity blob, open it with the customer's recovery
|
||||
// code R, and hand back EXACTLY ONE field: the offsite restic repository password.
|
||||
//
|
||||
// WHY ONLY ONE FIELD. The bundle also carries the Cloudflare tunnel token, the PBS access token and
|
||||
// the WG private key (see IdentityBundle). The caller in this flow — the in-guest controller, one
|
||||
// trust tier down — needs none of them, and returning them would widen the blast radius of a
|
||||
// controller compromise for no gain. Narrowing costs nothing here and is not recoverable later.
|
||||
//
|
||||
// WHY R NEVER TOUCHES DISK. `UnwrapIdentity` stages the BLOB and the recovered plaintext in a
|
||||
// `MkdirTemp` that it removes, and feeds R through the pty; R itself is never written. This wrapper
|
||||
// keeps that property: it takes R as an argument, passes it straight through, and holds no copy.
|
||||
// Callers must clear their own reference (the `R = ""` discipline in cmd/felhom-agent).
|
||||
//
|
||||
// The errors below are DISTINCT on purpose. "no blob", "wrong code" and "the blob predates the field"
|
||||
// are three different situations for the operator and only one of them is a fault.
|
||||
|
||||
var (
|
||||
// ErrNoEscrowBlob — the hub holds no sealed bundle for this host. Not a fault: no ceremony has run.
|
||||
ErrNoEscrowBlob = errors.New("escrow: the hub holds no sealed identity bundle for this host (no ceremony has run)")
|
||||
// ErrNoResticPassword — the bundle opened, but carries no repository password. Real and expected
|
||||
// for a pre-fork-4 blob (agent < v0.77.0, 2026-07-09): the field did not exist and CANNOT be
|
||||
// retro-fitted, because R is never retained. Distinguished from a wrong code so the operator is
|
||||
// not sent hunting for a mistyped recovery code that was typed correctly.
|
||||
ErrNoResticPassword = errors.New("escrow: the recovered bundle carries NO offsite repository password (a pre-fork-4 blob — the field did not exist when it was sealed and cannot be retro-fitted)")
|
||||
)
|
||||
|
||||
// BlobFetcher yields this host's own opaque identity-escrow blob. present=false is a clean "none".
|
||||
// An interface-free func field keeps this package free of any dependency on the hub client.
|
||||
type BlobFetcher func(ctx context.Context) (blob []byte, present bool, err error)
|
||||
|
||||
// OffsiteKeyRecoverer is the assembled links 6→8. Construct it with a fetcher; call it with R.
|
||||
type OffsiteKeyRecoverer struct {
|
||||
Fetch BlobFetcher
|
||||
}
|
||||
|
||||
// RecoverOffsiteRepoPassword fetches, unseals and extracts. It returns ONLY the repository password.
|
||||
//
|
||||
// A WRONG RECOVERY CODE FAILS CLOSED at the scrypt KDF inside UnwrapIdentity — `age -d` exits
|
||||
// non-zero and emits no plaintext, so there is no partial result and nothing is written anywhere.
|
||||
// That property is the crypto's, not a check here, which is why this function has no "validate R"
|
||||
// step to get wrong.
|
||||
//
|
||||
// NOTHING IS LOGGED BY THIS FUNCTION and no error it returns contains R, the password, or blob bytes.
|
||||
func (r OffsiteKeyRecoverer) RecoverOffsiteRepoPassword(ctx context.Context, recoveryCode string) (string, error) {
|
||||
if r.Fetch == nil {
|
||||
return "", fmt.Errorf("escrow: recoverer has no blob fetcher configured")
|
||||
}
|
||||
if recoveryCode == "" {
|
||||
return "", fmt.Errorf("escrow: the recovery code is required")
|
||||
}
|
||||
blob, present, err := r.Fetch(ctx)
|
||||
if err != nil {
|
||||
return "", fmt.Errorf("escrow: fetching the sealed bundle: %w", err) // carries no secret
|
||||
}
|
||||
if !present || len(blob) == 0 {
|
||||
return "", ErrNoEscrowBlob
|
||||
}
|
||||
bundle, err := UnwrapIdentityBundle(ctx, blob, recoveryCode)
|
||||
if err != nil {
|
||||
return "", err // already the fail-closed "the recovery code did not unwrap…" message; no secret in it
|
||||
}
|
||||
if bundle.ResticRepoPassword == "" {
|
||||
return "", ErrNoResticPassword
|
||||
}
|
||||
return bundle.ResticRepoPassword, nil
|
||||
}
|
||||
@@ -0,0 +1,189 @@
|
||||
package escrow
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// R-199 links 6→8, with REAL crypto (age is present on the build/demo host; ensureAge skips
|
||||
// elsewhere). These are the unit half of the session's question — "is the repository password
|
||||
// actually recoverable from the sealed bundle" — and the live half is the same equality on hardware.
|
||||
|
||||
const testR = "correct horse battery staple sedative anaconda wobbly kingdom placard yodel"
|
||||
|
||||
func sealBundle(t *testing.T, b IdentityBundle, r string) []byte {
|
||||
t.Helper()
|
||||
blob, err := WrapIdentityBundle(context.Background(), b, r)
|
||||
if err != nil {
|
||||
t.Fatalf("WrapIdentityBundle: %v", err)
|
||||
}
|
||||
return blob
|
||||
}
|
||||
|
||||
func fetcherFor(blob []byte) BlobFetcher {
|
||||
return func(context.Context) ([]byte, bool, error) { return blob, true, nil }
|
||||
}
|
||||
|
||||
// Scenario A (unit) — the recovered repository password is BYTE-IDENTICAL to the sealed one, and it
|
||||
// is the REPOSITORY password rather than some other field of a bundle that also parses.
|
||||
//
|
||||
// RED-PROOF: return bundle.PBSToken (or TunnelToken, or WGPrivateKey) instead of
|
||||
// bundle.ResticRepoPassword → a plausible-looking bundle yields a non-matching key → this FAILS.
|
||||
// That mutation is the shape of the bug that would otherwise ship silently, because every one of
|
||||
// those fields is a non-empty string that looks like a secret.
|
||||
func TestRecoverOffsiteRepoPassword_ReturnsTheRepositoryPassword(t *testing.T) {
|
||||
ensureAge(t)
|
||||
const repoPW = "0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef"
|
||||
blob := sealBundle(t, IdentityBundle{
|
||||
TunnelToken: "TUNNEL-TOKEN-NOT-THE-ANSWER",
|
||||
PBSToken: "PBS-TOKEN-NOT-THE-ANSWER",
|
||||
WGPrivateKey: "AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA=",
|
||||
ResticRepoPassword: repoPW,
|
||||
}, testR)
|
||||
|
||||
got, err := (OffsiteKeyRecoverer{Fetch: fetcherFor(blob)}).RecoverOffsiteRepoPassword(context.Background(), testR)
|
||||
if err != nil {
|
||||
t.Fatalf("recover: %v", err)
|
||||
}
|
||||
if got != repoPW {
|
||||
t.Fatalf("the recovered key is not the sealed repository password (len %d vs %d) — a different "+
|
||||
"field of the bundle was returned", len(got), len(repoPW))
|
||||
}
|
||||
// Belt: it must not be any of the OTHER fields, so a future refactor cannot satisfy the check
|
||||
// above by coincidence.
|
||||
for _, other := range []string{"TUNNEL-TOKEN-NOT-THE-ANSWER", "PBS-TOKEN-NOT-THE-ANSWER", "AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA="} {
|
||||
if got == other {
|
||||
t.Fatalf("the recoverer returned the wrong bundle field")
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Scenario B — a WRONG recovery code fails closed, the failure names no secret, and nothing is
|
||||
// written. The fail-closed property is the crypto's (age's scrypt KDF), which is why there is no
|
||||
// validation step here to get wrong — the test pins that it stays that way.
|
||||
func TestRecoverOffsiteRepoPassword_WrongCodeFailsClosed(t *testing.T) {
|
||||
ensureAge(t)
|
||||
const repoPW = "ffffffffffffffffffffffffffffffffffffffffffffffffffffffffffffffff"
|
||||
blob := sealBundle(t, IdentityBundle{TunnelToken: "t", PBSToken: "p", ResticRepoPassword: repoPW}, testR)
|
||||
|
||||
got, err := (OffsiteKeyRecoverer{Fetch: fetcherFor(blob)}).RecoverOffsiteRepoPassword(context.Background(), "not the recovery code at all")
|
||||
if err == nil {
|
||||
t.Fatal("a wrong recovery code MUST fail — a plausible-but-wrong bundle is the one outcome the design forbids")
|
||||
}
|
||||
if got != "" {
|
||||
t.Fatalf("a failed unseal returned %d bytes — there must be no partial result", len(got))
|
||||
}
|
||||
// The error may name the step; it may never name a secret.
|
||||
for _, secret := range []string{repoPW, testR, "not the recovery code at all"} {
|
||||
if strings.Contains(err.Error(), secret) {
|
||||
t.Fatalf("the failure message leaked a secret: %v", err)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A bundle with no repository password is its OWN answer, not a wrong-code error. Sealed before
|
||||
// fork-4 (agent < v0.77.0) the field did not exist; sending the operator to re-check a correctly
|
||||
// typed recovery code would be the wrong instruction.
|
||||
func TestRecoverOffsiteRepoPassword_PreForkFourBundle(t *testing.T) {
|
||||
ensureAge(t)
|
||||
blob := sealBundle(t, IdentityBundle{TunnelToken: "t", PBSToken: "p"}, testR)
|
||||
|
||||
_, err := (OffsiteKeyRecoverer{Fetch: fetcherFor(blob)}).RecoverOffsiteRepoPassword(context.Background(), testR)
|
||||
if !errors.Is(err, ErrNoResticPassword) {
|
||||
t.Fatalf("a pre-fork-4 bundle must report its own error, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// Scenario D at this layer — no blob is a clean, distinguishable answer.
|
||||
func TestRecoverOffsiteRepoPassword_NoBlob(t *testing.T) {
|
||||
rec := OffsiteKeyRecoverer{Fetch: func(context.Context) ([]byte, bool, error) { return nil, false, nil }}
|
||||
_, err := rec.RecoverOffsiteRepoPassword(context.Background(), testR)
|
||||
if !errors.Is(err, ErrNoEscrowBlob) {
|
||||
t.Fatalf("absent blob must yield ErrNoEscrowBlob, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// Scenario F — R persists NOWHERE. TMPDIR is redirected into the test's own directory, the unseal is
|
||||
// run for real, and the whole tree is then walked: no file may contain R (or the recovered password),
|
||||
// and the staging directory the unseal creates must be gone.
|
||||
//
|
||||
// RED-PROOF: write R to a temp file anywhere in the flow (e.g. add
|
||||
// `os.WriteFile(filepath.Join(work,"r"), []byte(recoveryCode), 0o600)` inside UnwrapIdentity before
|
||||
// its defer removes the dir — or simply drop that defer and let the plaintext staging survive) → the
|
||||
// walk finds it → this FAILS.
|
||||
func TestRecoverOffsiteRepoPassword_RLeavesNoTrace(t *testing.T) {
|
||||
ensureAge(t)
|
||||
const repoPW = "1111111111111111111111111111111111111111111111111111111111111111"
|
||||
tmp := t.TempDir()
|
||||
t.Setenv("TMPDIR", tmp) // os.MkdirTemp honours this — every staging dir lands under the walk
|
||||
|
||||
const wrongR = "wrong code entirely"
|
||||
blob := sealBundle(t, IdentityBundle{TunnelToken: "t", PBSToken: "p", ResticRepoPassword: repoPW}, testR)
|
||||
if _, err := (OffsiteKeyRecoverer{Fetch: fetcherFor(blob)}).RecoverOffsiteRepoPassword(context.Background(), testR); err != nil {
|
||||
t.Fatalf("recover: %v", err)
|
||||
}
|
||||
// A failed unseal must leave nothing either — exercise both paths before walking.
|
||||
_, _ = (OffsiteKeyRecoverer{Fetch: fetcherFor(blob)}).RecoverOffsiteRepoPassword(context.Background(), wrongR)
|
||||
|
||||
// THE PRIMARY ASSERTION IS EMPTINESS, not content. A content scan alone is defeatable by a later
|
||||
// call OVERWRITING the leaked file with a different secret — which is exactly how the first
|
||||
// version of this test passed its own red-proof while R sat on disk. Nothing in this test writes
|
||||
// under TMPDIR, so after both calls the tree must contain no files at all.
|
||||
var survivors []string
|
||||
err := filepath.Walk(tmp, func(path string, info os.FileInfo, err error) error {
|
||||
if err != nil || info == nil || info.IsDir() || path == tmp {
|
||||
return nil
|
||||
}
|
||||
survivors = append(survivors, strings.TrimPrefix(path, tmp))
|
||||
return nil
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if len(survivors) > 0 {
|
||||
t.Fatalf("the unseal left %d file(s) behind under TMPDIR: %v — R, the sealed blob and the "+
|
||||
"recovered plaintext all pass through there and none of them may outlive the call", len(survivors), survivors)
|
||||
}
|
||||
// Defence in depth: any secret that DOES appear anywhere is named, for every code used.
|
||||
_ = filepath.Walk(tmp, func(path string, info os.FileInfo, err error) error {
|
||||
if err != nil || info == nil || info.IsDir() {
|
||||
return nil
|
||||
}
|
||||
body, rerr := os.ReadFile(path)
|
||||
if rerr != nil {
|
||||
return nil
|
||||
}
|
||||
for label, secret := range map[string]string{"R": testR, "a wrong R": wrongR, "the repository password": repoPW} {
|
||||
if strings.Contains(string(body), secret) {
|
||||
t.Errorf("%s survived on disk at %s", label, path)
|
||||
}
|
||||
}
|
||||
return nil
|
||||
})
|
||||
// And the staging directories are gone, not merely free of secrets.
|
||||
entries, _ := os.ReadDir(tmp)
|
||||
for _, e := range entries {
|
||||
if e.IsDir() && strings.HasPrefix(e.Name(), "felhom-idesc-") {
|
||||
t.Fatalf("an unseal staging directory survived: %s", e.Name())
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A fetch failure surfaces as a fetch failure, not as a wrong-code error — the operator must not be
|
||||
// sent to re-read their recovery code because the hub was unreachable.
|
||||
func TestRecoverOffsiteRepoPassword_FetchErrorIsDistinct(t *testing.T) {
|
||||
rec := OffsiteKeyRecoverer{Fetch: func(context.Context) ([]byte, bool, error) {
|
||||
return nil, false, errors.New("hub: connection refused")
|
||||
}}
|
||||
_, err := rec.RecoverOffsiteRepoPassword(context.Background(), testR)
|
||||
if err == nil || !strings.Contains(err.Error(), "fetching the sealed bundle") {
|
||||
t.Fatalf("a fetch failure must say so, got %v", err)
|
||||
}
|
||||
if errors.Is(err, ErrNoEscrowBlob) || errors.Is(err, ErrNoResticPassword) {
|
||||
t.Fatal("a transport failure must not masquerade as a content verdict")
|
||||
}
|
||||
}
|
||||
@@ -307,3 +307,50 @@ func tail(b []byte, max int) string {
|
||||
}
|
||||
return s
|
||||
}
|
||||
|
||||
// IdentityEscrowResponse mirrors GET /api/v1/hosts/{host_id}/escrow (hub >= v0.94.0, R-199).
|
||||
// Present=false is a CLEAN answer, not a fault: the host simply has no sealed bundle yet.
|
||||
type IdentityEscrowResponse struct {
|
||||
HostID string `json:"host_id"`
|
||||
Present bool `json:"present"`
|
||||
IdentityEscrowB64 string `json:"identity_escrow_b64"`
|
||||
}
|
||||
|
||||
// FetchIdentityEscrow reads back THIS host's own opaque identity-escrow blob (R-199 link 6 — the
|
||||
// mirror of UploadEscrow, self-scoped server-side by the per-host key). The bytes are ciphertext: they
|
||||
// are useless without the customer's recovery code R, which neither the hub nor this agent ever holds.
|
||||
//
|
||||
// It is the ONLY retrieval this client performs, and it is deliberately narrow — no directive, no
|
||||
// K-escrow, no key rotation. The operator-driven DR path (recovery-mode re-enroll) is a different
|
||||
// endpoint with a different gate and is not reached from here.
|
||||
//
|
||||
// Errors are typed (transport vs HTTP) and never include the bearer token. The BLOB is never logged —
|
||||
// only its length.
|
||||
func (c *Client) FetchIdentityEscrow(ctx context.Context) (*IdentityEscrowResponse, error) {
|
||||
if c.hostID == "" {
|
||||
return nil, fmt.Errorf("hub: FetchIdentityEscrow requires a configured host_id")
|
||||
}
|
||||
url := c.baseURL + "/api/v1/hosts/" + c.hostID + "/escrow"
|
||||
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("hub: building escrow-fetch request: %w", err)
|
||||
}
|
||||
req.Header.Set("Authorization", "Bearer "+c.apiKey)
|
||||
req.Header.Set("Accept", "application/json")
|
||||
|
||||
resp, err := c.hc.Do(req)
|
||||
if err != nil {
|
||||
return nil, &TransportError{Err: err}
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
|
||||
raw, _ := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
|
||||
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
|
||||
return nil, &HTTPError{StatusCode: resp.StatusCode, BodyTail: tail(raw, 256)}
|
||||
}
|
||||
var out IdentityEscrowResponse
|
||||
if err := json.Unmarshal(raw, &out); err != nil {
|
||||
return nil, fmt.Errorf("hub: decoding escrow fetch: %w", err)
|
||||
}
|
||||
return &out, nil
|
||||
}
|
||||
|
||||
@@ -0,0 +1,102 @@
|
||||
package localapi
|
||||
|
||||
import (
|
||||
"context"
|
||||
"crypto/sha256"
|
||||
"encoding/hex"
|
||||
"errors"
|
||||
"net/http"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-agent/internal/escrow"
|
||||
)
|
||||
|
||||
// R-199 (agent v0.125.0) — the in-guest controller asks the agent to recover the offsite repository
|
||||
// password from the hub's sealed bundle, using the customer's recovery code R.
|
||||
//
|
||||
// WHY THE AGENT AND NOT THE CONTROLLER. Three reasons, all structural: the unsealing binary (`age`)
|
||||
// is an agent runtime dependency and is deliberately absent from the controller image; the sealed
|
||||
// blob is a HOST-scoped object whose only writer is this agent under the per-host key, so the read is
|
||||
// that write's mirror; and the controller is a trust tier down — it should receive one field, not a
|
||||
// bundle it has no use for.
|
||||
//
|
||||
// R'S HANDLING, WHICH IS THE TIGHTEST RULE IN THIS FLOW. R is the one secret in the system that
|
||||
// cannot be rotated, re-issued or recovered — it exists only in the customer's hands. Here it:
|
||||
// - arrives in the request body over the already-pinned local-API channel (the operator accepted
|
||||
// that crossing on 2026-08-04; the acceptance covers the CHANNEL, not carelessness at either end);
|
||||
// - is held in memory for the duration of one call and cleared on BOTH paths;
|
||||
// - is never written to disk, never an argument in a process list, and never logged at any level,
|
||||
// including inside an error;
|
||||
// - is never echoed: no response this endpoint can emit contains it.
|
||||
//
|
||||
// The request-level DEBUG middleware logs method/path/status/duration and never bodies — see
|
||||
// `logRequests`. Do not add a body dump.
|
||||
//
|
||||
// THE RESPONSE CARRIES THE PASSWORD AND ITS HASH. The hash is what this session's proof compares
|
||||
// (compare by hash, never by value). The password itself is present because the next link — placing a
|
||||
// recovered password so the existing repository opens — needs it, and building a hash-only seam now
|
||||
// would have to be torn out to add it. The controller's diagnostic reads only the hash.
|
||||
|
||||
type recoverOffsitePasswordRequest struct {
|
||||
VMID int `json:"vmid"`
|
||||
// RecoveryCode is the customer's R. NEVER logged, never persisted, never echoed.
|
||||
RecoveryCode string `json:"recovery_code"`
|
||||
}
|
||||
|
||||
// handleRecoverOffsitePassword fetches this host's sealed bundle, unseals it with R and returns only
|
||||
// the offsite repository password (plus its sha256, for hash-only comparison by the caller).
|
||||
func (s *Server) handleRecoverOffsitePassword(w http.ResponseWriter, r *http.Request, vmid int) {
|
||||
var req recoverOffsitePasswordRequest
|
||||
if !decodeBody(w, r, &req) {
|
||||
return
|
||||
}
|
||||
if !s.scopedFromBody(w, req.VMID, vmid, r.URL.Path) {
|
||||
return
|
||||
}
|
||||
R := strings.TrimSpace(req.RecoveryCode)
|
||||
req.RecoveryCode = "" // drop the decoded copy immediately
|
||||
if R == "" {
|
||||
writeErr(w, http.StatusBadRequest, "recovery_code is required")
|
||||
return
|
||||
}
|
||||
if s.escrowRecovery == nil {
|
||||
R = ""
|
||||
writeErr(w, http.StatusServiceUnavailable, "offsite key recovery is not configured on this agent (no hub client)")
|
||||
return
|
||||
}
|
||||
|
||||
ctx, cancel := context.WithTimeout(r.Context(), 60*time.Second)
|
||||
defer cancel()
|
||||
s.logger.Info("local-api: recovering the offsite repository password from the sealed escrow (R via body, never logged/persisted)", "vmid", vmid)
|
||||
|
||||
pw, err := s.escrowRecovery.RecoverOffsiteRepoPassword(ctx, R)
|
||||
R = "" // cleared on BOTH paths, before anything else can happen
|
||||
if err != nil {
|
||||
// Each situation gets its own status and its own words. None of them names a secret.
|
||||
switch {
|
||||
case errors.Is(err, escrow.ErrNoEscrowBlob):
|
||||
s.logger.Warn("local-api: offsite key recovery: the hub holds no sealed bundle for this host", "vmid", vmid)
|
||||
writeErr(w, http.StatusNotFound, "the hub holds no sealed recovery bundle for this host — no escrow ceremony has run")
|
||||
case errors.Is(err, escrow.ErrNoResticPassword):
|
||||
s.logger.Warn("local-api: offsite key recovery: the bundle opened but predates the repository-password field", "vmid", vmid)
|
||||
writeErr(w, http.StatusConflict, "the recovery code opened the bundle, but it carries NO offsite repository password (sealed before that field existed; it cannot be retro-fitted)")
|
||||
default:
|
||||
// Includes the fail-closed wrong-code case. The agent log records the STEP, never the code.
|
||||
s.logger.Warn("local-api: offsite key recovery FAILED (wrong recovery code, or the blob could not be fetched)", "vmid", vmid, "err", err)
|
||||
writeErr(w, http.StatusBadRequest, "the recovery code did not open the sealed bundle, or the bundle could not be fetched — nothing was written")
|
||||
}
|
||||
return
|
||||
}
|
||||
|
||||
sum := sha256.Sum256([]byte(strings.TrimSpace(pw)))
|
||||
// §8.6's lesson, applied: say exactly WHAT was recovered and what was NOT, so nobody reading this
|
||||
// concludes the wrong thing about the bundle's contents (which is how link 8 came to be missing).
|
||||
s.logger.Info("local-api: offsite repository password RECOVERED from the sealed escrow — returning that field ONLY "+
|
||||
"(the tunnel token, the PBS token and the WG key stay inside the agent and are not returned)",
|
||||
"vmid", vmid, "restic_pw_sha256", hex.EncodeToString(sum[:]))
|
||||
writeOK(w, map[string]any{
|
||||
"restic_repo_password": pw,
|
||||
"restic_pw_sha256": hex.EncodeToString(sum[:]),
|
||||
})
|
||||
}
|
||||
@@ -111,6 +111,14 @@ type HostMetricsProvider interface {
|
||||
}
|
||||
|
||||
// Options configures a Server.
|
||||
// EscrowRecoverer opens this host's sealed identity bundle with the customer recovery code and
|
||||
// returns ONLY the offsite restic repository password (R-199 links 6-8). An interface so the
|
||||
// localapi package needs no hub-client dependency and the route is testable without crypto.
|
||||
// R is an argument and is never retained by any implementation.
|
||||
type EscrowRecoverer interface {
|
||||
RecoverOffsiteRepoPassword(ctx context.Context, recoveryCode string) (string, error)
|
||||
}
|
||||
|
||||
type Options struct {
|
||||
ListenAddr string // bridge IP:port
|
||||
Cert tls.Certificate
|
||||
@@ -210,6 +218,11 @@ type Options struct {
|
||||
// GET /debug/logs. OPTIONAL — when nil the endpoint reports "not configured".
|
||||
LogRing *applog.Ring
|
||||
Logger *slog.Logger
|
||||
// EscrowRecovery (R-199, v0.125.0) is the offsite-key recovery seam behind
|
||||
// POST /escrow/recover-offsite-password. OPTIONAL — nil → that route reports "not configured"
|
||||
// (503) instead of failing obscurely. Satisfied by escrow.OffsiteKeyRecoverer.
|
||||
EscrowRecovery EscrowRecoverer
|
||||
|
||||
}
|
||||
|
||||
// defaultBackupCadence is the fallback /backup/due window when none is configured.
|
||||
@@ -273,6 +286,11 @@ type Server struct {
|
||||
netMountRoot string // the user-data namespace root for the network-mount role gate
|
||||
smbCredsDir string // where SMB creds files are written (out-of-band, 0600)
|
||||
escrowStagePath string // fork-4: 0600 staging file for the pushed restic repo password
|
||||
// escrowRecovery (R-199, v0.125.0) assembles chain links 6-8: fetch this host's own sealed
|
||||
// identity blob from the hub, unseal it with the customer's recovery code, return ONLY the
|
||||
// offsite repository password. OPTIONAL — nil (no hub client configured) makes
|
||||
// POST /escrow/recover-offsite-password answer 503 rather than pretending.
|
||||
escrowRecovery EscrowRecoverer
|
||||
intent IntentRecorder // slice 10 P3 (optional)
|
||||
guestBinds *GuestBindStore // F9 startup bind re-assert record (optional)
|
||||
formatJobs *FormatJobStore // F20-BUG3 detached-format job record (optional)
|
||||
@@ -423,6 +441,7 @@ func NewServer(o Options) (*Server, error) {
|
||||
netMountRoot: storage.NetworkMountRoot,
|
||||
smbCredsDir: o.SmbCredsDir,
|
||||
escrowStagePath: o.EscrowStagePath,
|
||||
escrowRecovery: o.EscrowRecovery,
|
||||
intent: o.Intent,
|
||||
guestBinds: o.GuestBinds,
|
||||
formatJobs: o.FormatJobs,
|
||||
@@ -518,6 +537,10 @@ func (s *Server) Handler() http.Handler {
|
||||
mux.HandleFunc("POST /escrow/stage-secret", s.withGuest(s.handleStageEscrowSecret))
|
||||
// fork-4 hygiene: wipe the staged secret once escrowed (controller calls this on confirm). Idempotent.
|
||||
mux.HandleFunc("DELETE /escrow/stage-secret", s.withGuest(s.handleWipeStagedEscrowSecret))
|
||||
// R-199 (v0.125.0): recover the offsite repository password from the hub-held sealed bundle,
|
||||
// using the customer recovery code supplied in the body. Returns that ONE field. See
|
||||
// escrow_recover.go for R's handling rules — they are the tightest in this package.
|
||||
mux.HandleFunc("POST /escrow/recover-offsite-password", s.withGuest(s.handleRecoverOffsitePassword))
|
||||
|
||||
// Controller-driven escrow ceremony (v0.88.0): preflight checklist, the detached root ceremony
|
||||
// job (fixed-argv sudo self-invocation), its status, and the ONE-SHOT in-memory R claim.
|
||||
|
||||
Reference in New Issue
Block a user