Files
felhom.eu/REPORT.md
T

22 KiB
Raw Blame History

REPORT — R-193/R-192 spike (no code) + R-195 phantom-customer alarm (hub v0.92.0)

Date: 2026-08-04 · Class: spike (Parts 03) + one fenced implementation (Part 4) Boundary held: the spike shipped no production code for R-193 or R-192, and Part 4 never touched the spike's subject matter.

Supersedes the earlier 2026-08-04 report (R-191 + R-190), per the overwrite convention. Nothing durable from it lived only here — it is recorded in hub/CHANGELOG.md, CONTEXT.md S-24/S-25 and OPEN-ITEMS.md R-190/R-191.


1. Confirmed baselines, re-read at the cited commit

Repo main @ commit on arrival Version End state
felhom.eu f456835bbcc4 ✔ matches the spec hub v0.91.1 (deployed felhom-hub:0.91.1) v0.92.0 built, pushed, manifest bumped, synced, live
felhom-controller 0887fd676d60 untouched (read-only)
felhom-agent 856a127 untouched (read-only)

Clean-tree gate passed before the build: git status --porcelain empty, HEAD == origin/main. Highest register ID on arrival was R-194; grepped before minting.


2. Q1Q8 — each answered or explicitly marked unmeasured

Full method, quotes and query output: documentation/audits/SPIKE-offsite-credential-recovery-2026-08-04.md.

Q2 and Q3 first — they decide whether the rest matters

Q2 — does a rebuilt controller reuse the escrowed repository password, or mint a new one? RULING: it MINTS A NEW ONE. Method — source: ConfigureOffboxApplyOffsiteTargetWriteOffboxSecrets, which generates a fresh 256-bit password whenever <DataDir>/offbox/repo_password is absent (felhom-controller/controller/internal/backup/offbox.go:392-401). A rebuilt guest has a fresh data dir. The only recovery path, InjectOffboxPassword, has exactly one caller in the entire repo (census: grep -rn "InjectOffboxPassword" --include=*.go . → 3 hits, 1 definition, 1 caller, 1 comment) — web/offbox_handlers.go:189, a web form a human pastes into. The apply-bridge never calls it. Method — live, without touching a box: host_escrow.restic_pw_sha256 vs host_escrow_superseded.restic_pw_sha256. demo-hp 8e03eddf…8a9e33aa…; demo-felhom 48741892…c60c8bc7…. Both CHANGED. This is bigger than R-193 as filed: demo-felhom, which R-193 records as having survived by luck, had 36 snapshots / 1,136,685,919 B before its rebuild and has reported snapshot_count: 0 in all 109 reports since. Its luck restored delivery, not the repository.

Q3 — what did the 04:15 run report? UNMEASURED, and the binary in the question is wrong. The run is daily at ~02:15 UTC (measured last_run values on both boxes). Neither box could run on 2026-08-04: demo-hp had no target at all until 07:15:47 UTC; demo-felhom's target was escrow_state: pending from 08-03 07:19:10 to 08-04 07:20:13, which OffboxRunnable blocks (configured AND escrowed). The decisive run is 2026-08-05 ~02:15 UTC. Source predicts a third outcome, neither 15 nor 1: same account + same repo path (u629488-sub3:/home/felhom-repo; repoPath is a compile-time constant) + new password + claimed: 1ensureOffboxRepo classifies orphanedErrOffboxOrphanedthe run refuses and shows the orphan card. Recorded as a prediction, not a measurement.

The rest

  • Q1 — what is one-shot. Full table in the spike. Headline: the hub's one-time provider password is one-shot but freely re-issuable; the box's SFTP key is regenerated on every apply by design; the restic repository password is the one nothing can restage — the agent's own source calls it "irreplaceable (unlike the SFTP access key…)" (felhom-agent/internal/escrow/identity.go:35-39). R-193 was filed against the recoverable secret; the damage is in the irreplaceable one.
  • Q4 — see §3 below.
  • Q5 — why one box recovered and the other not. CONFIRMED FROM THE LEDGER. demo-felhom's one_time_secrets row was created 2026-07-21 08:29:29 (traceable to a Re-issue at that exact timestamp) and sat UNCONSUMED for 13 days until the rebuild consumed it at 2026-08-03 07:17:58. Without one: exactly demo-hp — consume404 and a retry loop, 25 hours unprotected. Correction to the contrast: the recovery was of the plumbing only (Q2).
  • Q6 — why delivery is one-shot. The design's own reasons quoted from three sites. The security half, stated: the consume endpoint is gated by checkAuthCustomerthe customer's API key, the same credential as the config pull — so an automatic restage grants no new access to an unauthenticated party. What it genuinely widens is the exposure window: a live provider password becomes fetchable whenever the box reports no offsite target, rather than only in the seconds after an operator acts. Blast radius stays the customer's own label-scoped sub-account. The R-39(a) unconsumed-secret guard must survive any change.
  • Q7 — R-192's guard. BOTH HALVES CONFIRMED WITH NUMBERS. See §4.
  • Q8 — what each candidate costs. See §5.

3. Which of the two contradictory claims about Re-issue is true

The claim that Re-issue "only resets the transport credential" is TRUE about what ReissueCredentials does. R-39's "rotates the restic password" is FALSE.

From source (hub/internal/offsite/offsite.go:150-228), the function does four things: genPassword, ResetSubaccountPassword/ResetBoxPassword, SaveOneTimeSecret, and bookkeeping. There is no reference to a restic password anywhere in it, and none is possible — that password is generated on the box and never leaves it except into the R-wrapped escrow, which the hub cannot open.

But the bookkeeping is justified by the false premise, offsite.go:198-201: "the restic repo password just changed, so any existing key-escrow blob … is now STALE". The same wrong claim appears at api/handler.go:1067-1069 and in R-193's own row. Eighth entry in CLAUDE.md's table of comments asserting a guarantee the code does not provide — the first where the comment factually describes a different function. → R-196.

What corroborated it, and the corroboration is the strongest part. The task asked: if the repo password was untouched, what explains demo-hp's escrow going pending after the Re-issue? Fully explained, and not by the Re-issue: ApplyOffsiteTarget carries EscrowState over from the existing target, the rebuilt guest had none, so the pending fallthrough fired — and the box had just minted a new repo password, so pending was correct for a different reason. The proof they are independent is demo-felhom: it received no Re-issue at all (its only escrow_stale/offsite_reissued pair is dated 2026-07-21 08:29:29), its repo password changed anyway, and nothing marked its escrow stale for 13 hours. The mechanism is exactly inverted: escrow_stale is wired to the one path that does not change the repo password and absent from the path that does.


4. Every query used, with its fields and values

The DB snapshot was taken from the live pod and its freshness proven with a positive observable, not with "the query returned no error":

# Query / read Result
1 kubectl exec deploy/hub -- cat /data/{hub.db,hub.db-wal,hub.db-shm} The -wal is load-bearing: hub.db alone was 2 h 15 m stale (mtime 08:34 vs wal 10:45 CEST). Copying only the main file is the R-3 shape.
2 PRAGMA integrity_check ok
3 MAX(host_reports.received_at) vs datetime('now') 2026-08-04 08:45:38 vs 08:49:46newest row 4 min old
4 SELECT customer_id, status, created_at FROM customer_configs demo-felhom active · peti-felhom active · demo-hp active · drill-r50 blocked · david active (created 2026-08-01 16:51:49). sess-f is gone — confirmed, not assumed. drill-r50 untouched.
5 hosts / host_deletions / host_reports / reports counts for david 0 / 0 / 0 / 0 — never bound
6 events WHERE event_type='expected_dbdump_missed' AND created_at > '2026-07-25' only david, 03:00 UTC on 08-02, 08-03, 08-04
7 SELECT customer_id, MAX(received_at), COUNT(*) FROM reports GROUP BY customer_id demo-hp 1329 · demo-felhom 6690 · drill-r50 183 (last 07-29) · peti-felhom 482 (last 2026-07-15) — the reason peti-felhom is down and does not alarm
8 distinct $.offsite objects per box with first/last seen demo-hp: 15 snapshots / 40 927 044 B until 08-03 05:59:21, ABSENT 06:12:1908-04 07:12:01 (111 reports), then pending 07:15:47, escrowed 07:16:02, snapshot_count 0. demo-felhom: 36 snapshots / 1 136 685 919 B until 08-03 07:01:28, ABSENT ×1, then pending 07:19:1008-04 07:20:13, escrowed 07:20:28, snapshot_count 0
9 host_escrow + host_escrow_superseded restic_pw_sha256 the decisive measurement — both changed (§2)
10 one_time_secrets (no values read) demo-felhom created 2026-07-21 08:29:29, consumed 2026-08-03 07:17:58; demo-hp created+consumed 2026-08-04 07:11:51/07:12:06
11 $.claimed on the latest report of each box 1 on both — which selects ensureOffboxRepo's refuse branch
12 CountReportsOffsiteSince reproduced verbatim (ORDER BY id LIMIT 500) for demo-hp @ 2026-07-23 09:53:41 total=500, withOffsite=500, spanning 2026-07-23 09:53:472026-07-28 11:17:40
13 same window, no LIMIT total=1174, withOffsite=1063 (⇒ 111 without — matches #8)
14 full text of the live offsite_delivery_stuck message "…and 500 report(s) since carry no offbox target…" while OffsiteReportsSinceConsume was 500
15 COUNT(*) of offsite_credential_restaged 0, for every customer, ever
16 the shipped predicate run verbatim against live data (post-deploy) david0; demo-felhom/demo-hp/drill-r50/peti-felhom1

Instrument limits recorded rather than discovered later: one_time_secrets has customer_id as PRIMARY KEY and is last-write-wins by design, so it holds only the newest credential per customer — demo-hp's 2026-07-23 consume, on which R-192's whole diagnosis rests, was overwritten by yesterday's Re-issue and is no longer readable from that table. The 2026-07-23 09:53:41 anchor in query #12 comes from R-192's record, and that is stated in the spike rather than presented as a live read.


5. Ranked options, with a recommendation

rank option verdict
1 (d) the honesty pass — now Cheap, weakens nothing, and the only option that helps the incident already in flight. Two parts: correct R-192's message + log its refusal (R-196), and detect a changed restic_pw_sha256 — both hashes are already in the hub's DB and nothing reads them (R-197).
2 (c) agent-retained repo password — NOT PREVIOUSLY NAMED The only candidate aimed at the irreplaceable secret. The agent lives on the Proxmox host, survives a guest rebuild, already receives the password over the pinned local API and already writes it to a fixed 0600 path — it merely wipes it after the ceremony. Every seam exists. Price: a copy of the data key at rest on the host. The operator's trade to make.
3 (a) hub restages automatically Already exists as reissueOnReenroll's F3 leg, wired to the wrong event (behind handleHostEnroll's mint-once-reuse short-circuit; a guest rebuild leaves the hosts row intact). Worth doing after (c), never instead: on its own it restores the transport to a box whose data key is already gone, and would have made both boxes look healthy on 2026-08-04 while their snapshots were orphaned — strictly worse than the current loud failure.
4 (b) recover from escrow at re-bootstrap Not implementable as stated. The escrow is R-wrapped and zero-knowledge; the hub has no recovery code. "Recoverable at re-bootstrap" describes a customer-present ceremony — i.e. the manual form that already exists.

Recommendation: ship (d) now; then decide (c). Do not ship (a) first.

The operator's question — stated, and NOT answered

The irreplaceable secret is the restic repository password. It is generated on the box, destroyed by a guest rebuild, and the only surviving copy is sealed under the customer's recovery code. Every machine will be reinstalled. Do you want that key to survive a rebuild unattended — by keeping a copy on the Proxmox host, where you already have root — or do you want it to stay sealed and accept that every rebuild starts a new repository, provided the system says so loudly on the day it happens?

STOPPED here. Neither candidate was begun.


6. Part 4 — the phantom-customer alarm (SHIPPED)

Files changed

  • hub/internal/store/store.gonew (*Store).HasEverBoundHost(customerID) (bool, error): a live hosts row OR a host_deletions tombstone.
  • hub/internal/monitor/deadline.go — one guard at the top of CheckBackupDeadlines' customer loop, plus an unbound counter in the summary line.
  • hub/internal/monitor/deadline_unbound_test.gonew, 5 tests.

Commit: 7fff45d688fa (code + tests + spike + docs) · f581ac134964 (manifest bump).

The mechanism, established at source and corroborated live. The existing down-skip could never have covered this: it reads StalenessChecker.GetState(), whose map is seeded from store.GetCustomers()a query over the reports table (store.go:1055-1067). A customer with zero reports is in no row, gets no state, and GetState() returns "" rather than "down". The skip misses exactly the customer it would most obviously cover. Corroborated by query #7 and by the live post-deploy startup log — Staleness checker initialized: 2 ok, 0 stale, 1 down — where the 1 down is peti-felhom (host deleted 2026-07-15, 482 old reports) and david is absent from the census entirely.

The discriminator is "was a host EVER bound", not "has a report arrived", which is the point: a box that was installed, bound and went silent has a hosts row and keeps alarming. Fail-open on a read error — an unreadable binding must never suppress a real alarm. The deferral is logged with its own counter (the v0.73.0 Part-7 precedent). The anchored-verdict structure is untouched, and david's record was not modified.

Sibling verdicts checked rather than assumed. The census of every customer-iterating checker: deadline.go backup half (already guarded by reportJSON == ""), restoretest.go (guarded by latest == ""), staleness.go + offsite.go (iterate GetCustomers(), report-derived — a hostless customer never appears), offsite_delivery.go (a hostless customer derives DeliveryNoSecret and is skipped). The DB-dump half was the only verdict that could fire; the guard covers both verdicts in that function uniformly.

Test results — the suite run and the exit code read BEFORE any commit, never in one command.

go build ./...  → BUILD OK
go vet ./...    → VET OK
go test ./...   → rc=0, 17 packages ok, no FAIL lines

The 13 TestCheckBackupDeadlines*/TestHasEverBoundHost tests were run with -v and the === RUN lines were checked, per CLAUDE.md's rule that a -run pattern matching nothing prints ok and exits 0.

THE RED-PROOF THAT MATTERS. TestCheckBackupDeadlines_BoundButNeverReported_StillAlarms is the case the change could break — a real box that was installed and never phoned home, indistinguishable from david on every signal except the one the guard reads. With the guard removed it still passes, which is the correct result and proves the test is aimed at the guard's risk, not its presence. The companion red-proof of the fix itself, observed:

=== RUN   TestCheckBackupDeadlines_NeverBoundHost_Silent
    deadline_unbound_test.go:58: a customer with NO host ever bound must raise NOTHING;
                                 got [expected_dbdump_missed]
--- FAIL: TestCheckBackupDeadlines_NeverBoundHost_Silent

— verbatim the event david sent three mornings running. deadline.go restored from a backup taken before the mutation; full suite re-run green afterwards.

Deployed version. felhom-hub:0.92.0 built + pushed (sha256:ca71140cf54a…), manifests/hub.yaml bumped 0.91.1 → 0.92.0 in git, ArgoCD hard-refresh + deliberate sync (never kubectl set image). Verified: app Synced / Healthy, rollout statusdeployment "hub" successfully rolled out (rc=0), running image gitea.dooplex.hu/admin/felhom-hub:0.92.0, pod hub-6775889d95-gc8sk 1/1 Running, clean startup log.

Is the 05:00 alarm confirmed silent? NO — CONFIRMATION IS PENDING, and it is stated rather than assumed. The hub's own log says deadline-check: next run at 2026-08-05 05:00 CEST (in 17h53m57s); there is no manual trigger (CheckBackupDeadlines has exactly one caller, the scheduler at cmd/hub/main.go:598). An absent alarm tomorrow would in any case be weak evidence, so a positive observable was taken instead: the shipped predicate was run verbatim against a fresh copy of the live DB (query #16) — david0 (will be skipped), every other customer → 1 (behaviour unchanged). The 05:00 firing itself is still owed.


7. OPEN-ITEMS.md rows, and the capability map

  • R-192 — updated with Q7's confirmed numbers; NOT closed, and now carries a new reason not to fix it in isolation (a successful auto-restage would have protected the plumbing while the data was already gone).
  • R-193 — updated with the whole spike; NOT closed, state moved to WAITING-ON-OPERATOR, owner CC + operator. The old "15+ vs 1 decides it" line is struck through and replaced with the measured answer plus the third predicted outcome.
  • R-195 (new) — the phantom-customer alarm. SHIPPED, hub v0.92.0.
  • R-196 (new)escrow_stale wired to the path that does not change the repo password, absent from the one that does. OPEN.
  • R-197 (new) — the hub holds both restic_pw_sha256 values and reads neither. OPEN.

The capability map DID need touching, and the §N.5 escalation did NOT fire — the row was not overclaiming. 00-capability-map.md's offsite row already recorded reinstall-continuity: "a recreated data volume that orphaned the repo … is now CLASSIFIED … + a move-aside (never-delete) reset", shipped controller v0.142.0 and live-fired 2026-07-18. That independently corroborates the Q3 prediction from a different direction. What it did not record — and now does — is that this fires on an ordinary, planned, unattended guest rebuild on every box, with the measured 2026-08-03/04 instance, the hash evidence, and the inverted escrow_stale wiring. Status unchanged (PROVEN-LIVE): the guard is correct and is not the gap; the cadence is.

ROADMAP.md's newest row is R-171 and none of R-190…R-197 has one, so there was nothing to collapse. One correction was warranted under the coupling rule: R-1's row claimed the v0.57.0 reinstall arc makes offsite re-issue "auto-fire on re-enrollment" — scope-corrected in place, since that means host re-enrollment after a host delete, and a guest rebuild never reaches it.

CONTEXT.md gained S-26 (the spike's Q1/Q2 headlines + the three corrections that must not be re-inherited) and S-27 (the R-195 ruling and its mechanism). REUSE.md gained the HasEverBoundHost row with its "never a substitute for a liveness check" trap. STATUS.md rewritten for the operator, one screen.


8. CI, task ids, and --no-verify

commit what CI task id run # conclusion
7fff45d688fa R-195 code + tests + the spike + all docs 126 72 success
f581ac134964 manifests/hub.yaml 0.91.1 → 0.92.0 127 73 success

Matched by head_sha via curl -s "https://gitea.dooplex.hu/api/v1/repos/admin/felhom.eu/actions/tasks?limit=5" — a PULL check, not a reliance on the failure mail. --no-verify was NOT used. The pre-push hook ran and reported gates OK - push proceeding; python3 scripts/repo_gates.py was additionally run in full (all five gates OK) before committing.

Teardown: the spike provisioned nothing — no guest, no VM, no storage, no credential. The only artefact is a hub-DB snapshot in the session scratchpad, outside the repo.

Read-only discipline held: no ReissueCredentials, no ceremony, no credential rotation, no change to any box's offsite configuration, no snapshot deleted, no customer record modified. drill-r50 untouched; sess-f's removal confirmed rather than assumed.


9. Observations — noticed and NOT acted on

  1. The orphaned ciphertext is retained and keeps consuming quota, with nothing to prune it. resetOrphanedRepo is move-aside-never-delete by design (correct), but 40.9 MB + 1.14 GB of unopenable data will sit against a 50 GB soft quota indefinitely. S-24's ep0 prune jobs cover the PBS namespaces, a different tier. Not filed — it belongs to whichever option is chosen for R-193.
  2. host_escrow_superseded holds only two rows in the entire database, both from 2026-08-04. That is why the spike could not establish whether demo-felhom's 2026-07-21 Re-issue re-sealed an unchanged password — the predicted false-staleness shape of R-196. Retention policy for that table was not investigated.
  3. Both demo guests were rebuilt on consecutive mornings (demo-hp 08-03 06:12, demo-felhom 08-03 07:17, both to controller 0.192.0 with new config hashes). Why is outside this spike's scope and was not chased.
  4. api/handler.go:1067-1069's comment is the third copy of R-196's false premise, and it sits in a different package from the other two. Whoever fixes R-196 should grep rather than fix the one they were shown.
  5. peti-felhom is active with its host deleted since 2026-07-15, judged only by the down skip. It is correct today and would start alarming the moment it reported once. Not a defect, but the shape is worth knowing.
  6. The -shm/-wal trap generalises. Any future session reading the hub DB from the pod must copy all three files. Reading hub.db alone today would have produced confident answers from a file 2 h 15 m stale — the R-3 shape with a different filename.