store.New opened the DB with `?_journal_mode=WAL&_busy_timeout=5000`, which is mattn/go-sqlite3 syntax. The driver is modernc.org/sqlite, whose applyQueryParams reads only _pragma/_time_format/_time_integer_format/_txlock/_inttotime and IGNORES anything else WITHOUT AN ERROR. So the hub ran in rollback-journal mode with busy_timeout=0 for its entire life while its own source said otherwise. Surfaced as a false HOST STALE banner: in rollback-journal mode a reader excludes a writer, so rendering an operator page blocks a host report; the hub 500s, the agent waits its full 15-minute interval without retrying, and staleness fires at 30 minutes — two collisions is a false alarm plus an operator email. 13 collisions in one pod lifetime; the alarm fired twice on 2026-08-02 for a host that was up two days and reconciling throughout. The observable that proved it: a 128 MB /data/hub.db with no -wal/-shm beside it while the DB was open. Fix: ?_pragma=journal_mode(WAL)&_pragma=busy_timeout(5000)&_txlock=immediate. _txlock=immediate is not optional — database/sql's Begin() is DEFERRED, so a read-then-write tx must upgrade its lock and a failed upgrade is SQLITE_BUSY_SNAPSHOT, which busy_timeout does NOT retry; this store has 10+ db.Begin() sites and they are all write paths. Every test asserts what the DATABASE reports, never the DSN string — a string test would have passed for the whole life of the bug. Red-proof: restoring the shipped DSN reproduces journal_mode="delete", the missing -wal, and the live "database is locked (5) (SQLITE_BUSY)". Operational consequence handled: a WAL DB cannot be copied by taking hub.db alone — a bare `cat` opens cleanly and silently omits the newest writes. The break-glass retrieval in operations/nodes.md used exactly that; it and the recovery-inventory note are now WAL-aware.
261 KiB
v0.88.0 — the WAL that never was (2026-08-02, R-172)
The hub has never actually been in WAL mode. store.New opened the database with
?_journal_mode=WAL&_busy_timeout=5000 — mattn/go-sqlite3 syntax — while the driver is
modernc.org/sqlite, whose applyQueryParams reads only _pragma, _time_format,
_time_integer_format, _txlock and _inttotime. Everything else is ignored without an error.
So the hub ran in the default rollback-journal mode with busy_timeout=0 for its entire life, while
its own source said otherwise — a configuration asserting an invariant the code did not provide.
How it surfaced. A false HOST STALE banner for demo-felhom-8363b5 while the agent was up two
days and reconciling normally. In rollback-journal mode a reader excludes a writer, so rendering an
operator page can block a host report; the hub then returns HTTP 500, the agent logs
hub: report failed; keeping current interval and waits its full 15-minute interval, and
staleness fires at 30 minutes. Two consecutive collisions = a false alarm + an operator e-mail.
Measured: 13 SQLITE_BUSY collisions in one pod lifetime, and the alarm fired twice that day
(19:12:32 and 20:42:32 CEST) for a host that was never down.
The observable that proved it: a 128 MB /data/hub.db with no -wal/-shm file beside it
while the database was open. In WAL mode those files must exist.
The fix is one DSN, and each parameter earns its place:
?_pragma=journal_mode(WAL)&_pragma=busy_timeout(5000)&_txlock=immediate
journal_mode(WAL)— readers and one writer proceed concurrently, so a page render can no longer block a report. It is a property of the database FILE, so it persists once set.busy_timeout(5000)— writers still serialise; without a timeout SQLite returnsSQLITE_BUSYimmediately rather than waiting._txlock=immediate— the one that is easy to miss, and WAL + busy_timeout alone would not cover it.database/sql'sBegin()is DEFERRED, so a transaction that reads then writes must upgrade its lock, and a failed upgrade isSQLITE_BUSY_SNAPSHOT, whichbusy_timeoutdoes not retry. This store has 10+db.Begin()sites and they are all write paths (customer delete/reset, wg, appliance, pbsdr, telemetry, log bundles). Without this the fix would leave a known un-retryable path open.
Every test asserts what the DATABASE reports, never the DSN string — a test on the string would
have passed happily for the entire life of the bug. Five tests: the runtime pragma values; the
-wal/-shm files existing beside an open DB (the production signature, pinned); a reader not
blocking a writer (the consequence, not the mechanism); concurrent writers waiting instead of
erroring; and racing read-then-write transactions. Plus TestSQLiteDriverIgnoresMattnStyleParams, a
guard on the ROOT CAUSE: it fails if someone "tidies" the pragmas back to the familiar mattn form,
and skips itself with instructions if a future driver starts honouring them.
Red-proof: restoring the shipped DSN reproduces the live failure exactly — journal_mode = "delete",
the -wal absent, and a write FAILED while a read was open: database is locked (5) (SQLITE_BUSY).
Operational consequence, handled rather than discovered later: a WAL database cannot be copied by
taking hub.db alone — a committed transaction may still be in hub.db-wal, so a bare cat yields
a copy that opens cleanly and silently omits the newest writes. That is the worst shape for a
credential lookup, and the break-glass retrieval in documentation/operations/nodes.md used exactly
that command. Both it and the _recovery-inventory note are now WAL-aware (copy the -wal, shred
both).
Retries (options b and c in R-172) were NOT added. With readers no longer blocking writers and
the upgrade path covered, a SQLITE_BUSY reaching an HTTP handler should now be rare enough to be a
real signal. If any appear after this, they mean something else and a retry would hide it. Revisit
only on evidence.
v0.87.0 — the Setup tab stops claiming a host-install version it cannot know (2026-08-02)
R-94, all three legs, closed by deletion rather than derivation. The customer page's Setup Command card read "Day-0 host bootstrap for host-install 1.19.0". The served script was 1.22.0, and had been since 14 July — nineteen days of an operator-facing number that was simply wrong, with a version number's authority behind it.
Why deriving it is not achievable honestly. The Option-1 command downloads
felhom-host-install.sh from the website at run time, and the website git-syncs main every
thirty seconds (R-110). The hub therefore cannot know which version a given box will run — not at
build time, not at render time. Any literal there is a guess. The comment that guarded the old const
already half-admitted this ("Display-only… the served script is always current"). R-94(a) offered
derive it, or delete it; deleting removes the drift class permanently instead of automating it.
What changed.
internal/web/configs.go:const hostInstallVersion, thepageData.ScriptVersionfield and its assignment are gone. A NOTE stands in their place recording why there is deliberately no constant here, so the next person does not helpfully re-add one.internal/web/templates/customer_unified.html: the sentence now says the command always fetches the current installer and renders no version at all.internal/web/render_test.go: the assertionstrings.Contains(html, hostInstallVersion)compared the constant to itself and passed at any value — demonstrated green with the const set to9.9.9while the served script was 1.22.0. Deleted, not replaced: there is no longer a version to assert. Thedata-customer-idand static-fallback assertions stay.scripts/hostinstall_gates.pygate 1 inverts: it used to require the hub const to EQUALSCRIPT_VERSION; it now asserts the hub carries no host-install version literal at all, matched in six code shapes across every.go/.htmlunderhub/. Comments are deliberately not stripped — a//inside a URL string literal would truncate the scan and blind the gate — so the patterns match declarations, fields, assignments and the template action, never prose.scripts/felhom-host-install.sh: comment only,SCRIPT_VERSIONuntouched. It claimed the gate keeps the hub copy equal, an invariant that no longer exists; a comment asserting an invariant the code does not provide is a wish.
Red-proofs. Restoring the const fails the rewritten gate 1 on three of its six shapes. The old
render_test.go assertion passes with the const at 9.9.9.
Live verification: endpoint-level (no browser on DooPlex) — the customer Setup tab is fetched and grepped for a version literal.
v0.86.0 — Copy works without revealing, and every copy branch reports itself (2026-07-31)
Found by the operator, in the way that matters: it cost a real login. The v0.84.0 Console access
card shipped its Copy button disabled until a Reveal. Clicking it did nothing, silently — so the
clipboard kept whatever was already in it, which was another host's console password from an
earlier reveal. That got pasted into demo-hp's PVE login, which failed with no explanation. The box
logged a plain password check failed for user (root); the credential was never at fault, and there
was nothing on screen to say the copy had not happened.
A copy button that silently no-ops is worse than no copy button, because the operator has no way to distinguish "copied" from "did nothing" — and the stale value it leaves behind is a valid secret for a different machine, so the resulting failure looks like a stale-credential problem and sends you diagnosing the wrong thing.
Copy now works without revealing — and that is the safer default, not a concession. The secret goes straight to the clipboard and never renders on screen, so it cannot be shoulder-surfed or caught in a screenshot. Reveal is still there for when you need to read it (typing at a console).
Three silent-failure branches closed, all in the same eight-line function:
| Branch | Was | Now |
|---|---|---|
| Not yet revealed | button disabled, click = no-op |
fetches and copies |
navigator.clipboard absent (insecure context) |
if (navigator.clipboard) → silently skipped |
shows the password instead and says why |
writeText() promise REJECTED (permission / no user gesture) |
promise ignored — the operator believes it copied | shows the password instead and reports the refusal |
The success path now names the host: "✓ Copied demo-hp-bb76ea's root@pam password to the clipboard." The clipboard is fleet-wide and every box has a different console password, so "copied" alone cannot say copied for which box — precisely the confusion that produced the incident.
One retrieval path, shared. fetchConsolePassword is used by both buttons, and the endpoint is
defined once (the data-reveal-url attribute) and read back with getAttribute, so Copy cannot drift
onto a different — unaudited — URL than Reveal. A test asserts the URL appears exactly once.
Server-side nothing changed: both buttons hit the same CSRF-gated endpoint and both write the same
recovery_credential_revealed event, which is correct — the register records accesses, and a copy
is an access.
Tests 566 → 568, both pinning this regression: the Copy button must not ship disabled, and every
outcome branch must carry a message. Red-proof: re-adding disabled reproduces the shipped bug and
turns the first test red.
v0.85.0 — Network card: a host's addresses are visible at last (2026-07-31)
Pairs with agent v0.119.0 and is useless without it — the agent is what reports the addresses.
A managed box's LAN IP was not shown anywhere in the hub, because nothing reported it. The host
report carried no address of any kind. The only IP reachable from the UI at all was the WireGuard
one, on /offsite's peer table keyed by pubkey — so an operator could go peer→host and never
host→peer, which is the direction anyone actually asks in.
The host page grows a Network card: every routable address the box holds, one row per
(interface, address), plus a WireGuard row. On demo-felhom that is vmbr0 192.168.0.162/24 and
tailscale0 100.70.170.35/32 — with the PVE web console reachable at
https://<the LAN address>:8006, which is the thing the operator wanted and could not get.
WireGuard is rendered as TWO facts, deliberately. WGAssignedIP is the hub's own allocation
(wg_peers — desired state, authoritative) and WGConfirmed is whether the box reports actually
holding it. Showing the allocation alone would make a peer that was never applied look healthy —
the same shape as reading a timestamp that records an attempt as if it recorded a result. A
mismatch renders not confirmed by the box; there is a test for exactly that case, and a red-proof
that pins it (hard-wiring WGConfirmed = true turns it red).
The split is keyed on the ALLOCATION, not on the interface name. wg-felhom is the agent's
current unit name; a UI keyed on that string would silently mis-render the day it changes. Comparing
the reported address against the hub's allocated one uses the identity that survives a rename.
An old agent renders UNKNOWN, never "no addresses". Below agent 0.119.0 the field is absent
from the wire, and an absent signal is not a negative result — the page says "this host's agent does
not report its addresses — they are unknown, not absent" and names the version needed. Rendering an
empty list there would have stated something false about the host. Red-proofed: deleting the branch
makes the page claim the host has no routable address.
No new store table and no new ingest path — the report is already stored opaquely, and
GetWGPeerForHost already existed with no UI consumer. This is parse + render.
Files: hub/internal/web/hosts.go (parseHostAddresses, hostNetworkView, hostNetwork,
hostDetailData), hub/internal/web/templates/host_detail_body.html,
hub/internal/api/testdata/host-report.golden.json (the cross-repo contract, moved in lockstep with
the agent's copy).
Tests 559 → 566; four red-proofs (the inert view-model, unconditional confirmation, the old-agent
branch, and the report fixture being the REAL wire from --selftest=hub) each run, observed failing,
and reverted.
v0.84.0 — Break-glass console credential on the host page (2026-07-31)
The credential existed and was not reachable when it was wanted. Every Felhom-installed box has
had a strong random root@pam console password since TASK G1 — set on the box by
felhom-host-install.sh step 4b, vaulted in the hub at day 0, live for three hosts today, and used
for real during the sshd incident. The only way to read it back was a hand-written curl against
/api/v1/admin/hosts/<id>/recovery-credential carrying the global operator key — a different
secret from the hub login password, kept out-of-band. In practice the PVE web console on a demo box
felt locked.
The host page grows a Console access card. By default it states only that a credential is
vaulted, for which user, and when it was last set. A Reveal button fetches the plaintext on
demand and shows it for 60 s with a Copy button; masking clears the JS variable, and the mask also
fires on a second click and on visibilitychange → hidden. A host with nothing vaulted says so, and
says why (byo host, or step 4b never ran), with no Reveal control at all.
The secret is never rendered into the page — that constraint shapes the whole change. The render
path calls a new store.GetHostRecoveryMeta, whose struct and whose SELECT both omit the secret
column, so it is structurally incapable of carrying it; hostDetailData gains exactly three keys
(RecoveryVaulted, RecoveryUsername, RecoverySetAt). The plaintext crosses the wire only in the
response to POST /hosts/{id}/reveal-recovery-credential — Cache-Control: no-store, CSRF-gated at
the ServeHTTP level, POST precisely so that gate applies and so a secret is never retrievable by
URL alone (prefetch, history, referrer). The load-bearing test asserts the canary appears nowhere
in the rendered response — attribute, comment, inline script or JSON blob.
Deliberately NOT the customer_unified.html data-secret widget, which embeds the plaintext in
the page HTML on every load: acceptable for one customer's retrieval passphrase, not for console root
on every box in the fleet (it survives in the bfcache, in "save page as", and in any DOM-capturing
screenshot). That widget is untouched and recorded as an observation.
Transparency, matching the log-pull precedent. A delivered reveal writes one
recovery_credential_revealed event (info, source hub, Hungarian) on the host's customer timeline
— SaveEvent alone, no dispatcher call, so nobody is emailed. Two reveals write two events: the
register records accesses, not states. A 404 is not an access and writes nothing. An unbound host
reveals fine and writes no event (no customer to tell); the [INFO] hub line is then the only
record, and it carries the username and a length — never the password.
The global-key API path is untouched, by design. handleAdminGetRecoveryCredential is the
break-glass route for when the hub UI is the thing that is broken; coupling it to the session layer
would remove exactly the independence that makes it a fallback.
Recorded as a real trade, not a free one: the hub session password alone now unlocks console root
on every managed box, where retrieval previously also required the global API key. Accepted for a
single-operator, HU-geo-fenced hub that already stores these passwords in plaintext at rest — and the
plaintext-at-rest half is now filed as R-133 (envelope-encrypt host_recovery.secret under a KEK
held outside the DB, so a hub DB backup stops being a fleet-wide console-credential dump).
Files: hub/internal/store/host_recovery.go (+GetHostRecoveryMeta, HostRecoveryMeta),
hub/internal/web/hosts.go (+handleHostRevealRecoveryCredential, hostDetailData),
hub/internal/web/server.go (route, above the bare /hosts/ catch-all),
hub/internal/web/templates/host_detail_body.html (card + fetch-on-demand script).
Tests 550 → 559; four red-proofs (A page-leak, B audit event, D CSRF gate, E route order) each run,
observed failing, and reverted. The E proof is a seam test driving ServeHTTP: a handler-level test
cannot see that defect, because the handler is correct and simply never runs.
v0.83.0 — R-109 + R-122: the recipe assembly stops dropping sections (2026-07-30)
Pairs with agent v0.118.0 (R-106 + R-109). The agent half is useless without this one.
AssembleDRRecipe's two shape structs are ALLOW-LISTS, and nobody had noticed. The doc comment sold
hostHalfShape/appHalfShape as forward-compat — "encoding/json drops any unknown top-level key" — which
is true and is also the trap: a section an emitter adds is silently discarded until it is named in both
the shape struct and AssembledRecipe. No error, no log, no failing test. The section is simply not in the
file the operator downloads.
R-122 (found this session) — that already happened, and it shipped. The controller has emitted
offsite_restic since fork-4 — the offsite restic repo's non-secret coordinates, whose entire purpose is
"so DR knows WHERE to recover from". The hub stored it intact for every real customer
(peti-felhom, demo-felhom, demo-hp all carry it in dr_recipe.app_half_json today) and appHalfShape
never listed the key, so no delivered recipe has ever contained it. Verified both ways before the fix: the
stored half has it, GET /customers/demo-felhom/dr-recipe.json did not.
R-109 — and it would have happened again the same day. The agent's new backup_target (which storage
holds the local whole-guest archives) is a new top-level host-half section. Without this commit it would
have been stored and dropped exactly like offsite_restic, and the R-109 fix would have read as shipped
while changing nothing an operator can see.
Both keys are now on hostHalfShape / appHalfShape / AssembledRecipe, and the allow-list comment says
what it actually is, plus the rule: adding a recipe section is a TWO-REPO change.
Tests: 3 new, all consequence-level and all built on the halves production really stores — read verbatim
out of the hub's own dr_recipe table (the pre-existing drHostHalf/drAppHalf constants are hand-written
and OMITTED offsite_restic, which is precisely why the drop stayed green for the feature's whole life).
TestAssembleDRRecipe_CarriesEveryEmittedSection enumerates every section both emitters produce and fails
on any that does not survive assembly — the guard the allow-list needed and never had.
..._NamesTheLiveBackupTargetAmongTwoCandidates asserts the delivered recipe names felhom-backup at
/mnt/hdd_1 and not the frozen local, and refuses to run if the fixture stops posing that problem.
..._UnknownBackupTargetSurvivesVerbatim pins that the agent's explicit unknown reaches the operator AS an
unknown and does not acquire a storage_id on the way through.
Red-proofs: 2, each mutation asserted to have landed before running — drop offsite_restic from the
allow-list (the R-122 defect restored) → 2 tests fail; drop backup_target → 4 fail, naming the section.
go build + go vet rc=0; suite rc=0, 17 packages, 0 FAIL.
internal/store/dr_recipe.go—backup_target+offsite_resticon both the shape structs andAssembledRecipe; the allow-list warning.internal/store/testdata/dr-recipe.golden.json— both new sections +namespace_state.internal/api/testdata/host-report.golden.json— synced byte-identical with the agent's copy (f4bc3554…).
v0.82.0 — R-120: the vouch path refuses a golden the fleet has already outrun (2026-07-30)
The mechanism half of R-120. The golden's version is the controller it bakes
(felhom-agent configs/build-golden.sh:345 defaults GOLDEN_VERSION to ${CONTROLLER_IMAGE##*:}), so a
golden left behind the newest deployed controller means every fresh install lands on stale
application code. On the R-120 occurrence that stale code shipped a customer-facing falsehood: a box
installed from the 0.185.1 golden told a customer whose backup drive had fallen out that "the backup is
on the same disk as the system" — false, the drive was gone — and offered a different drive as the
remedy. 0.186.0 is the release that made that message true, and no new box had it.
Why a gate and not a reminder. This gap has opened three times — R-111 (the golden's agent 17
releases behind), R-115 (an agent built and deployed but never published), R-120 (this). The first
two were closed by re-baking and remembering; remembering then failed again. And R-29 is the standing
proof that a check nobody runs is worse than none, because it reads as coverage:
hostinstall_gates.py sat RED and invoked by nothing across three version bumps while every report said
green, and hub_confirm_gate.py has never run at all.
So the distinguishing property is not does a check exist but does it block:
- It lives in
handleSetArtifacts(internal/web/configs.go), immediately before the only write — the sole UI path tostore.SetArtifactManifest. It therefore runs on every vouch without anyone choosing to run it. A script inscripts/asserting the same fact would have been a fourth orphan. - It REFUSES (operator ruling, 2026-07-30), with an operator-legible flash naming the remedy, rather than warning.
- Signal:
store.NewestReportedControllerVersion()— the highest controller version any box has reported, fromreports.controller_version(the columnSaveReportdenormalises). Semver-compared in Go, notMAX()in SQL, which would rank 0.99.0 above 0.186.0 — a pair this fleet has actually shipped. No outbound call, no new credential.
Fail-open in exactly two cases, both deliberate: an empty golden field (clearing the manifest is a legitimate act) and an unknown fleet version (a new hub must be able to vouch its first golden).
Known blind spot, stated rather than papered over: a controller no box has ever run is invisible to this signal, so a golden baked behind an unreleased controller still passes. That is a real limit, and it is not the failure that has bitten — all three instances were "deployed newer than baked".
A near-miss worth recording. The first draft read guests.controller_version — a column that exists
in the schema (store.go:294) and that nothing writes. That gate would always have seen "" and
failed open: inert, i.e. precisely the R-29 shape it exists to prevent. Caught by grepping for a writer
before trusting the column.
Tests: 4, through the production handler over httptest, never an injected seam — because a gate
that can be inert is the thing this gate exists to prevent, and three shipped defects in this project
were fully green with the seam disconnected. Refusal asserts both the flash and that the manifest
was not written (a gate that redirects and saves anyway reads as enforcement while providing none);
plus the allow cases, both fail-open cases, and the semver-ordering case. Red-proof: deleting the block
makes the stale golden vouchable and both refusal assertions fail.
v0.81.0 — E-2: the absent backup target gets its own signal (2026-07-29)
Hub half of E-2, and it ships FIRST by necessity: an event type the hub does not allowlist makes
POST /event return 400 and the event vanishes (the R-97a failure). The controller cannot emit
backup_target_absent until this is live.
E-2's Phase 0 established that an absent backup target has no prompt signal today. The drive-gate
path stops apps and logs a WARN but emits nothing — NotifyStorageDisconnected is defined and never
called anywhere in the controller (verified against the gitignored-cmd/ trap with a positive
control). A drive that is only a backup target has no apps to stop, so it is entirely silent. The
sole signal is the tier's own failure at its next due cycle, i.e. up to ~24 h on the daily local
tier — the R-100 shape, where a real fault is visible only after a deadline elapses.
Added to BOTH registers, because each half fails differently and the second failure is the quiet one:
allowedEventTypes(internal/api/handler.go) — without it the event is lost at the door;customerMessages(internal/notify/templates.go) — without it the event IS delivered, but the customer receives the controller's raw operator English instead of Hungarian, and nothing looks broken.
backup_target_absent is deliberately not folded into storage_disconnected. That one says "a
drive went away and some apps may have stopped"; this one says "the thing that makes your backup
survive a disk failure is gone" — a different customer action and a different operator urgency.
Hungarian copy states the CONSEQUENCE, not just the fact:
„A rendszermentés meghajtója nem érhető el — amíg vissza nem csatlakoztatod, a teljes rendszermentés nem készül el."
backup_target_restored is the paired recovery (info severity — the existing recovery pattern;
severityNotifies is untouched and NOT widened).
Tests + red-proofs. api.TestBackupTargetEventTypesAreAllowlisted and
notify.TestBackupTargetCustomerMessagesArePresent pin the pair; a third test pins that the copy
names what is at risk and what happens, so a future shortening to a bare „Meghajtó hiányzik." cannot
pass. All three red-proofed with the mutation VERIFIED to have landed first — the initial attempt
silently no-op'd (gofmt had realigned the map to three spaces) and the test "passed", which would
have been a false proof.
v0.80.0 — R-100: staleness counts from the last SUCCESS (2026-07-28)
OffsiteChecker.isStale counted from last_run, which the controller writes unconditionally at the
end of every run including failures. It therefore asked "how long since we last TRIED" — so a tier
failing on every single night refreshed the clock nightly and read as perfectly fresh forever. It now
counts from last_success (controller v0.181.0).
What the defect is NOT, corrected after checking. The old comment here said "a recent-but-failing
run is NOT stale (backup_failed owns that signal)", and that was true — backup_failed does fire
for a failing offsite run, nightly, and reaches the operator (live hub DB: 5 operator sends). The real
defect is defeated defence in depth: this checker is the hub-side, pull-based net that exists to be
independent of controller-pushed events, and anchoring it on a field the failing controller keeps
refreshing made it depend on the very thing it backs up. F-HUB — this campaign's own finding, the hub
dropping an event under SQLITE_BUSY with no retry — is exactly that loss.
Three branches, each deliberate:
- never ran (no
last_run) — unchanged v0.73.0 anchored behaviour. Still keyed onlast_run, notlast_success, on purpose:last_runanswers "has anything ever happened here", and a box whose first run failed has alast_runand nolast_success— that is a run, not a newborn. - legacy (
last_runset, nolast_success) — degrades explicitly to the oldlast_runbehaviour, logged once per customer. Treating absence as failure would alarm the whole un-upgraded fleet at once; treating it as success keeps the bug. Same degrade direction as R-88 Part 2'sage_state. - anchored — counts from
last_success.last_statusis deliberately not consulted: "error ⇒ stale" pages on every transient blip, which is the F-A1 noise path. One bad night is tolerated because the threshold simply keeps running from the last good run.runningis a real wire value (a report captured mid-run) and is likewise not a verdict.
The alarm text had to change with the verdict. emitStale still said last run 8h ago while firing
on a six-day-old success — a true alarm that reads as a false one. staleAge now separates the two
diagnoses: "runs are happening and failing — check the error, not the schedule" versus "the offsite
leg is silently not running". last_success joins the event details.
Fixtures are the real wire shapes from 4000 live reports (ok ×2269, absent ×541 always with an
empty last_run, error ×27, running ×7), not invented JSON.
Red-proofs, all observed failing: restore the last_run anchor → a tier that has not succeeded in 6 days reads as FRESH; delete the never-ran branch → a newborn box alarmed; collapse to
last_status == "error" → a single transient failure alarmed; delete the legacy degrade → a legacy controller alarmed — that is a fleet-wide alarm storm on an un-upgraded fleet.
Felhom Hub — Changelog
v0.79.0 — R-97c: make the operator-only claim TRUE (2026-07-27)
v0.78.0 shipped a comment asserting that whole_guest_backup_failed / _recovered were operator-only
because they have no customerMessages entry, "so the dispatcher structurally cannot route them
to a customer". That was false, and the code says so plainly:
templates.gotreats a missing entry as a fallback to the raw message, not a block —hunMessage := customerMessages[eventType]; if hunMessage == "" { hunMessage = message };- the only customer gate is
isEventEnabled(prefs.EnabledEvents, ...)— configuration.
So a customer with whole_guest_backup_failed in their enabled list and an email set would have been
sent the raw English operator text about a backup they can take no action on. Proven by running the
new test against the v0.78.0 shape: it emails customer@example.com.
This is the EffectiveProtected shape — a doc comment claiming a property the code stopped
providing, which is how the samba false alarm survived.
The fix: an explicit operatorOnlyEvents register, checked at the top of processCustomer
before prefs are consulted, so no customer configuration can opt in. The skip is logged
(status=skipped, error_message=operator_only, channel=customer) rather than dropped — a silent
drop is indistinguishable from a delivery that never happened.
Deliberately not implemented as "a missing customerMessages entry blocks delivery": several
types rely on the raw-message fallback on purpose (offbox_enlarge_blocked's dynamic Hungarian text
is customer-grade and a template would discard its numbers), so turning the fallback into a gate
would change behaviour well outside this concern.
The recovery type is listed too, even though its customer leg is pairing-gated on a "sent" row that cannot exist — relying on that would make one type's safety a consequence of another type's routing, true today and silently untrue the moment the failed event became customer-visible.
The handler.go comment now states the actual mechanism and warns that allowlisting a type does not
make it operator-only.
Tests +4, all run under the breaking configuration (customer has the event enabled AND an email), not today's safe one. 17 packages ok.
v0.78.0 — R-97a: the whole-guest backup tier gets a voice (operator-only) (2026-07-27)
internal/quiesce had no route to the hub at all. On 2026-07-27 three failed whole-guest backups and
twelve app-stack stop/starts produced zero events. This is the hub half of the fix.
Two new operator-tier event types, allowlisted with no customerMessages entry:
whole_guest_backup_failed (error) and whole_guest_backup_recovered (info).
Deliberately NOT backup_failed/backup_completed. Both of those carry customer-facing Hungarian
templates AND sit in demo-felhom's live enabled_events — reusing them would have emailed the
customer „A biztonsági mentés sikertelen" while the backup was still retrying behind the R-88
breaker. A customer can take no action on a failed whole-guest backup. Same pattern as R-85's
restore-test types: in the allowlist so the chain works, out of customerMessages so the dispatcher
structurally cannot route them to a customer.
The recovery rides the F11 pairing branch. whole_guest_backup_recovered is severity info, and
severityNotifies drops info — routing it normally would store the event and never mail it, so the
operator would be told a tier broke and never told it healed. Adding it to recoveredPairedDownTypes
puts it on the recovery branch, which runs BEFORE the severity gate. Its customer leg needs no special
handling: that leg is pairing-gated on a customer-channel "sent" row for the down type, and the down
type can never produce one.
Per-tier operator cooldown (cooldownTierSuffix). The operator cooldown was keyed
customerID + ":" + eventType — correct for an event describing ONE thing, wrong for one describing
ONE TIER when a box has two: felhom-pbs failing at 09:00 would swallow local failing at 09:20 for
the rest of the hour. The key now appends the tier from the event details when present, so no
existing event type's cooldown behaviour changes. Widening it for everything would turn one hourly
app_start_failed into one per app — a flood, not a fix.
Tests +8. build/vet/test rc=0, 17 packages ok.
v0.77.0 — R-85 Part 2: a restore-test result becomes a SIGNAL (2026-07-26)
Until now a failed restore-test was a [WARN] line in the ingest handler and nothing else — no
event, no notification, no gauge. That was true for the local tier that was already being tested,
so the loudest DR signal this system produces was, in practice, inaudible. Rotating tiers (agent
v0.104.0) without this would only mean two tiers can fail silently instead of one.
Two signals, deliberately NOT merged
| event | meaning | severity |
|---|---|---|
restore_test_failed |
a run completed and did not pass — something is broken NOW | error |
restore_test_stale |
a tier has not been proven within its interval — nothing has necessarily broken; we no longer know | warning |
Merging them would collapse "your DR is broken" into "your DR is unverified", and the second is the one that quietly becomes the first. The staleness wording says "unverified, not known-broken" and a test asserts that phrasing.
Anchored, per R-81 — not re-derived
A tier never proven on a newborn box is UNKNOWN, not FAILED, until an anchored window elapses. This monitor family has made the opposite mistake three times (hub v0.12.0, v0.73.0, R-81); this is a NEW monitor written straight after the third, so it copies R-81's verdict structure and boundary-test discipline rather than inventing a fourth shape. The deferral is logged once — a quiet check must never be indistinguishable from one that did not run.
restoreProvenStaleAfter = 7d is derived, not guessed: a 24 h cadence rotating oldest-first across
two tiers proves each about every 2 days, so 7 days tolerates ~3 consecutive missed opportunities
before alarming — and sits comfortably inside the 2-week offsite retention, so a tier is never called
stale against an archive that is about to be pruned.
How per-tier proof is recovered
The agent reports only its latest restore-test and its store is in-memory, so the latest report alone cannot answer "when was the OTHER tier last proven?". The hub's retained host-report window can — the same R-81 mechanism, reused rather than re-solved with a wire change.
Also
- Both types registered in
allowedEventTypes. They are hub-generated, but that map is the project's single register of legitimate event types, and R-77's lesson was that a type missing from it ships as an inert seam. - Operator-tier only — neither has a
customerMessagesentry, so the dispatcher cannot route it to a customer. A customer can take no action on a failed restore-test, and "a visszaállítási teszt nem sikerült" would frighten without informing. A persistently unproven DR tier may eventually warrant a customer-visible statement; that needs copy review, not a side effect of this task. - A tier the box does not HAVE is never reported stale (the Slice-C gate, same reasoning).
Fixed — a time bomb introduced in Slice C
TestCheckBackupDeadlines_RestartBlindWindow_NoEvent hard-coded the literal incident timestamp
2026-07-18T18:31:06Z while comparing against the REAL clock. Harmless while one 26 h threshold
covered every tier; once Slice C gave the offsite tier an 8-day limit it became a bomb — the test
passed all day on 2026-07-26 and began failing at 18:31 UTC, exactly 8 days after that instant.
Now relative. A test that passes at commit time and fails hours later is worse than one that fails
immediately, because it lands on whoever is next in the file.
Tests
+10, full suite green (17 packages, rc=0, vet unpiped). Red-proofs observed:
- B — log-and-stop (the pre-R-85 shape) yields
a FAILED restore-test must EMIT an operator event; got 0 event(s). The assertion is that a NOTIFICATION IS EMITTED; the hollow version checks for a log line, which passes against exactly the code this replaces. - D — removing the anchor yields
a newborn box must NOT alarm; got restore_test_stale: local tier: NEVER successfully restore-proven in 0s of watching.
v0.76.0 — R-82 Slice C: tier-aware backup thresholds (2026-07-26)
R-81 merged every backup signal into one "newest" and judged it against a single 26 h limit. That
was right while a box had exactly one whole-guest tier. backupStaleAfter's own comment recorded
why it stops being right:
"The moment PBS moves to a WEEKLY cadence, a perfectly healthy weekly snapshot is >26h old six days in seven and this constant alarms on it."
Each tier is now judged against its own threshold. R-81's structure is preserved intact — three-valued verdicts, absence anchored at first contact, one distinct reason per failure mode — and its boundary test is untouched.
Added
offsiteBackupStaleAfter = 8d(7-day cadence + a day of headroom — the same ratio 26 h gives a 24 h cadence).backupStaleAfterkeeps 26 h and now names the host tier only.splitTiers— classifies a report's evidence into host and offsite tiers.assessTier— R-81's exact logic, parameterised by tier and threshold.newestBackupEvidenceByTier— the retained-window scan, per tier.
The Slice-A.4 rule, implemented
A PBS-targeted vzdump appears in both backups[] (as a Backup with target_id:"felhom-pbs")
and pbs_snapshots[] (enumerated independently from PBS). Classification is therefore by TARGET
TYPE — target_id → storage_targets[].name → .type == "pbs" — never by array membership.
Getting that wrong would let a PBS backup make a stale host tier look fresh, silently losing the
daily tier's alarm. Pinned by TestTierAware_PBSTargetedVzdumpIsNotHostEvidence.
storage_targets is used rather than pbs_dr.storage_id because the latter is null on a box that
has a PBS storage but no DR descriptor yet (drill-r50 was exactly that shape).
A tier is only judged when the box HAS it
expected gates each tier on a configured storage of that kind, or evidence for it. Without that,
every box without an offsite tier would alarm as soon as the anchor elapsed — the
absence-is-not-failure mistake R-81 exists to prevent, re-introduced one level down. When NEITHER
tier is identifiable (an old agent reporting no storage_targets and no target_id) the pre-Slice-C
combined path runs unchanged, so nothing regresses on a fleet mid-upgrade.
Changed behaviour (intended)
A 30 h-old offsite snapshot no longer alarms — under a weekly tier it is healthy. Three existing fixtures asserted the old merged threshold; each still asserts an alarm, now at the correct limit (9 days for offsite, 30 h for host). No assertion was weakened to make the code pass.
⚠️ Recorded limitation — the hub infers cadence from storage TYPE
"PBS ⇒ weekly" is an inference, not a fact the box tells us. defaultBackupTarget is "felhom-pbs",
so a box that never sets local_backup_target would run PBS as its daily primary tier and the
hub would judge it against 8 days — seven days of blindness. No box is in that shape today (both
demo boxes set local, and the installer pins it), but it is a latent mis-classification of exactly
the kind that became R-80. The real fix is the agent reporting each tier's actual cadence in the
host-report; own task.
Tests
Full suite green (17 packages). Red-proof observed: giving tierOffsite the host threshold — i.e.
restoring the merged limit — fails the 6-day case with
offsite tier: newest backup is 144h0m0s old (limit 26h0m0s), verbatim the cry-wolf this slice
removes. Restored.
Replayed against the live hub DB through the real store queries:
demo-felhom host=07-26T14:38Z offsite=07-26T12:21Z -> OK
demo-hp host=07-26T07:06Z offsite=none -> UNKNOWN (offsite watched 119h of a 192h grace)
drill-r50 host=none offsite=not expected -> MISSED (host tier, no evidence in 29h)
No customer email would be sent by this deploy. demo-felhom is clean; demo-hp defers correctly and will alarm in ~3 days if its offsite tier stays empty (the true R-82 finding, arriving on schedule); drill-r50's alarm is a true positive and it has no customer channel.
v0.75.0 — R-81: "no signal" is not "bad signal" — the backup deadline check is ANCHORED (2026-07-26)
The third instance of one bug class, fixed as a class. expected_backup_missed fired on
demo-felhom, demo-hp and drill-r50 simultaneously at 03:00 UTC, and the demo-felhom one reached the
customer channel claiming newest backup is 176h0m0s old. Nothing was wrong: three vzdump
archives were on disk (07-24, 07-25, 07-26). Full evidence:
documentation/audits/DIAG-backup-missed-2026-07-26.md.
Cause. The agent's backup record store is IN-MEMORY
(felhom-agent/internal/backup/store.go — "lost on restart; the cadence re-populates"). The R-50
island migration restarted the fleet at 12:44 UTC; the next backup landed at 07:03 the following
morning. In between, every host-report carried backups: [], and assessBackupFreshness read empty
as no backup exists. For demo-felhom it then fell through to the only surviving evidence — a PBS
snapshot from 07-18 — and reported its age as the customer's backup age.
The class. hub v0.12.0 (expected_backup_missed daily for every healthy customer — looked for
an event nobody emits), hub v0.73.0 (offsite_stale minutes after a healthy repair — never-ran
branch had no time anchor), and now this. All three: absence of signal treated as evidence of
failure. The invariant is now written at the head of assessBackupFreshness with all three
instances named, and pinned by a boundary test whose name says what it protects.
Changed
assessBackupFreshnessreturns a three-valued verdict —verdictOK/verdictUnknown/verdictMissed, replacingmissed bool. Absence is UNKNOWN, not a fault. It becomes a fault only once it outlives an anchored window. Still pure (nowand the evidence are injected) — that purity is why the incident was diagnosable and why this fix is provable.- The check now reads hub HISTORY, not just the latest report. New
store.GetHostReportsSince+monitor.newestBackupEvidenceanswer "when did I last SEE evidence of a backup?" across a bounded 7-day lookback (backupEvidenceLookback). The agent's store is point-in-time and forgets across a restart; the hub's retained reports (90 d) do not. This is the whole fix for the 07-26 shape — no agent change, no new persisted state, and semantically exactly the right question. The scan stops at the first sufficiently-fresh evidence, so the healthy path reads one row; only the genuinely-broken path walks the lookback. - The absence anchor is first contact — new
store.GetFirstHostReportAt. Absence is graded against how long the hub has been watching, reusing the existingbackupStaleAfter(26 h) as the grace exactly as v0.73.0 reused offsitestaleAfter. No new knob. A zero anchor fails toward visibility (the v0.73.0 legacy-shape precedent). CheckBackupDeadlineslogs the deferral. A deferred UNKNOWN emits one INFO naming the reason, and the summary line gained abackup unknown (deferred)counter — so a quiet check is never indistinguishable from a check that did not run (v0.73.0 Part-7 precedent). At most one line per customer per day.- Reason strings split, not collapsed. Absence-over-time, unanchored absence, deferred-newborn, stale-timestamp, failed-verify and unparseable each keep a distinct message. The entire 07-26 diagnosis turned on reading the exact string; a test enforces distinctness.
NOT changed (deliberate)
backupStaleAfterstays 26 h, and no tier-awareness was built. ⚠️ Landmine recorded in the constant's comment: it applies to whichever tier is newest, so once PBS moves to a weekly cadence a healthy weekly snapshot is >26 h old six days in seven and this will alarm on it. Per-tier thresholds cannot be built before the per-tier cadence config exists — R-82 owns both halves. Building it now would be speculative generality.parseBackupTimeuntouched — the agent emits clean RFC3339Zand the parse branch is not implicated. Its silentcontinueon an unparseable timestamp is a latent member of this same class and is recorded as an observation only.- The DB-dump half of
CheckBackupDeadlinesis event-based and correct — untouched. - The customer-facing Hungarian copy (
notify/templates.go:106) is untouched here. The DIAG found it overstates scope (it reads as all backups failed, but this check only covers the host/PBS tier); that copy change was not in this task's scope. - The agent's in-memory
Storeis the cause; making the host-report truthful rather than merely defensively interpreted is R-84.
Tests
508 total (was 493), +15 in internal/monitor/deadline_anchor_test.go. Companion red-proofs observed
and restored for all three acceptance scenarios:
- A (restart blind window must not alarm) — removing the history fold-in reproduces
newest backup is 176h0m0s old (limit 26h0m0s), verbatim the message demo-felhom actually sent. - B (a genuinely dead box must still alarm) — the naive "absence is always silent" fix fails
TestBackupFreshness_NoEvidenceBeyondAnchor_Alarmsand two contract rows. This is the test that makes A safe: a suite proving only A would pass against an implementation that never alarms. - C (a fresh box is not born failing) — the literal pre-fix branch fails four contract rows plus the end-to-end newborn case.
Replayed against the real thing: the actual host-reports the hub held at 2026-07-26 03:00 UTC
(600 / 417 / 77 retained rows) fed through the new policy → demo-felhom OK (window evidence
2026-07-25T06:30:14Z), demo-hp OK (2026-07-25T10:23:31Z), drill-r50 UNKNOWN (newborn,
watched 17 h < 26 h grace). All three silent.
v0.74.0 — allow local_api_endpoint_drift (controller v0.173.0 / R-77) (2026-07-26)
One line in allowedEventTypes. It is not optional: handleEvent 400s an unknown event_type
("Invalid event_type"), so the controller's new drift alert would have been silently inert without
it — the exact seam-wiring failure class this project has hit four times. Shipped with the controller
that emits it, not after.
Operator-only, error severity (drift never self-heals), and deliberately not an agent_channel_*
type: during the 2026-07-25 island-migration outage the generic "agent unreachable" event was the only
signal for 17.5 h and it hid a specific, fixable config fault. Naming the cause separately from the
symptom is the whole point. No customer notification toggle, matching the other agent_channel_* and
host_* operator events.
Source: documentation/audits/DIAG-agent-channel-2026-07-26.md.
v0.73.2 — sync hostInstallVersion → 1.19.0 (R-50 island host-install) (2026-07-25)
hostInstallVersion (the script version the operator customer page's install-command generator
advertises) bumped 1.16.0 → 1.19.0 to match felhom-host-install v1.19.0 (R-50 island default). The
hostinstall_gates.py F-1 gate requires the two move together; this also clears the pre-existing
1.16.0↔1.18.0 drift. Render test (render_test.go) confirms the value reaches the page. No behaviour
change beyond the advertised version string.
v0.73.1 — allowlist disk_health_degraded (controller v0.169.0 disk-health) (2026-07-24)
Adds disk_health_degraded to allowedEventTypes so the controller's per-disk SMART degradation
notification (v0.169.0) is ingested and delivered instead of 400-rejected. Like offbox_enlarge_blocked,
there is deliberately no customerMessages entry — the controller sends a dynamic Hungarian message
(disk label + the triggering attribute names), which the templates.go fallback preserves; a static entry
would discard the specifics. Test TestHandleEvent_DiskHealthDegradedAccepted (red-proof: drop the
allowlist line → 400).
v0.73.0 — offsite_stale no longer cries wolf on a newborn tier (ISO-train v1.25.0 Part 7) (2026-07-23)
Origin (operator, 2026-07-23 12:01 CEST): demo-hp's offsite was repaired and escrowed at 10:01Z and
offsite_stale fired MINUTES later (last_run:"" … threshold 48h0m0s) — the never-ran branch had
no time anchor, so "enabled + escrowed + never ran" was instantly stale on the first tick, with
remedy copy ("check the controller/schedule") that was wrong for the moment's true state.
The fix (no new constant, one boundary) — monitor.OffsiteChecker.isStale:
- Boundary ruling (recorded): never-ran staleness is owned by
offsite_staleONLY in the v0.72.0applieddelivery state; pre-applied never-ran shapes belong tooffsite_delivery_stuckalone — one state, one owner, never both. Structurally this was already true (Checknil-skips reports without the offsite object, and a report CARRYING the object ISappliedby the v0.72.0 definition) — now it is pinned by an explicit boundary test that also asserts the delivery checker remains that shape's only voice. - Anchor: never-ran staleness =
appliedAND (now − anchor) > the EXISTING 48 h threshold, where anchor = the newest ofone_time_secrets.consumed_at(delivery completed; via the v0.72.0GetOneTimeSecretInfo) and the customer's escrow-blob timestamp (host_escrow.updated_at/created_atvia the newLatestEscrowTimeForCustomer— reduced in Go, not SQL MAX, because the two timestamp formats would misorder lexicographically). Runs become possible only at the ceremony, so the ceremony anchors the clock. No separate grace knob: the existing threshold, anchored properly, IS the grace. - Anchor-less legacy shape (no secret row, no escrow row): pre-fix behavior kept — stale on sight, fail toward visibility. Ran-before behavior byte-untouched (explicit both-direction tests: old run + fresh anchor still fires; fresh run + old anchor stays silent).
- One INFO log line on the FIRST observation of a deferred newborn (
never-ran within the anchored threshold (anchor …)) — the anchored evaluation is provable live without per-sweep spam.
Red-proof: never-ran branch reverted to the pre-fix return true → the fresh-anchor fixture and
the escrow-anchor fixture both fired offsite_stale:warning (FAIL observed) → fix restored.
Origin: documentation/audits/DIAG-f10-demo-hp-offsite-2026-07-23.md — demo-hp sat 2 days with the
customer card claiming "Provisioned … delivered to the controller once" (static copy) while the box
had NOTHING: the day-0 managed update killed the apply-bridge after password-consume, and the hub —
holding both signals (one_time_secrets.consumed_at, 153 offbox-less reports) — read neither.
One detector, four consumers:
- Detector (
internal/offsite/delivery.go,DeliveryStateFor): per-customerOffsiteDeliveryStatefrom the secret row × report offsite-presence —applied(precedence: the box's own report wins) /consumed_awaiting_apply(the burned-credential shape) /staged_awaiting_consume/no_secret. The applied+stale-staged edge (demo-felhom's live shape: key-auth-first never consumes) staysappliedPLUS a visible stale-staged flag. New store reads:GetOneTimeSecretInfo(timestamps only, value never selected),LatestReportOffsitePresence,CountReportsOffsiteSince,LastEventAt, + a test-only timestamp back-dater (the PBSDR seam pattern). - Customer card (
config_form_body.html+deliveryViewFor): the static "delivered once" claim is GONE; the card renders the derived state with its age (badge idiomn-ok/n-warn/ n-neutral; consumed goes amber past 30 min; the stale-staged info line names the specimen's age). Render test per branch (the v0.70.1 template-gate rule). - Loud event:
offsite_delivery_stuck(WARNING → operator email; severity gate untouched — explicit tests pin that info would be silent) when consumed_awaiting_apply persists ≥ 1 h; per-customer 24 h cooldown, DURABLE viaLastEventAtover the events table (a hub restart neither floods nor resets). - Self-heal (R-71c):
OffsiteDeliveryChecker(shared 60 s ticker) invokes the EXISTING Re-issue path (web.Server.ReissueOffsiteForCustomerbehind the narrowmonitor.OffsiteReissuerinterface — the pbsdrheal precedent; armed only when the provisioner is configured) when the shape is unambiguous: consumed ≥ 1 h + ≥ 4 consecutive offbox-less reports since consume + zero offbox evidence. One restage per customer per 24 h (durable); every firing surfaces asoffsite_credential_restaged(WARNING → operator) — repeats are repeating events, never a silent retry loop. The R-39(a) guard:SaveOneTimeSecretclobbers by design (Re-issue depends on supersede), so the heal RE-READS the secret row immediately before acting and refuses unless it is still a consumed row — the TOCTOU (operator Re-issue landing mid-tick) is the shape the guard kills. Red-proofs run and recorded (guard removed → the operator's fresh secret gets clobbered, test fails onreissue calls = 1; both rate limits removed → duplicate fire).
Companion: felhom-controller v0.161.0 (the box-side truthful empty state). The self-heal ships unit-proven + red-proofed, NOT live-fired — no broken box exists and none was broken for it; it arms on the next natural occurrence (or a staged drill). R-71(a) (day-0 ordering) stays open — separate spec.
v0.71.0 — paired recovery mails, prefs seeding at claim, priority headers, operator test leg (2026-07-22)
Origin: documentation/audits/AUDIT-power-outage-recovery-2026-07-22.md F11 (recovery is silent),
F12 (prefs row optional → customer never notified), F14-light (delivered ≠ noticed). Live proof of
the gap: the demo customer got „A szerver nem elérhető!" at 15:29 and was never told it recovered.
- Paired recovery notifications (F11) —
dispatcher.goprocessRecovery, an explicit eventType branch inProcessEventBEFORE the severity gate (severityNotifiesand the checkers'emitTransitionseverities are byte-untouched;*_recoveredstaysinfo). Operator always gets both edges (existing 1 h per-type cooldown); the customer gets recovery iff the customer was mailed the paired stale/down — pairing evidence isstore.LastCustomerSentAtovernotification_log(customer channel, status=sent,node_recovered→{node_stale,node_down},host_recovered→{host_stale,host_down}), ties resolve to no-mail (flap-safe).enabled_eventsis deliberately NOT consulted for recovery. Suppressions log at INFO with the reason.FormatOperatorEmailrenders ✅ for*_recovered;customerMessagesgainshost_recovered. - Prefs seeding at claim (F12) —
claim.Engine.MarkClaimedseedscustomer_notificationsfrom the registeredcustomer_configs.emailon the unclaimed→claimed transition via newstore.SeedNotificationPrefs(INSERT OR IGNORE — never touches an existing row; empty email = no-op; a seed failure never fails the claim). Default set (critical-only, Viktor may adjust): node_down, backup_failed, disk_critical, host_disk_critical, storage_fill_critical, offbox_repo_orphaned. - Empty-email no-clobber guard (F12) —
handleSavePreferences: a push with an empty email preserves a stored non-empty address (events + cooldown still apply); a non-empty push updates everything. Phase-0a fact: controller 0.160.0 guards both push legs itself (cmd/controller/main.go:821startup skips empty email;web/handlers.go:1532refuses empty-with-events), so the clobber was latent — this is the hub-side belt for older/rogue boxes. - Priority headers (F14-light) —
sendEmailFn/sendEmailgain aheadersparam; Resend payload carries"headers"only when non-empty.priorityHeaders(severity): error/critical →X-Priority: 1+Importance: high; warning/info → none. Mechanism probed live pre-implementation (Resend accepted, HTTP 200, mail id34d3f7f3…). - Operator test leg — the
testevent now also mails the operator (✅ <id>: teszt / operator channel OK, priority headers forced) — one click proves customer channel + operator channel + header rendering. Fixed a latent nil-deref found here:sendTestEmaildereferencedprefs.EmailwhileGetNotificationPrefsreturns(nil, nil)for a customer with no row — a test event for such a customer (e.g. demo-hp) panicked the dispatcher goroutine. - Tests: 17 new (449 → 466) across store/notify/claim/api; 4 red-proofs run + reverted (pairing
removed, upsert-seed, guard removed, unconditional headers) — see
REPORT.md.
v0.70.1 — the ghost customer's Delete button must exist (2026-07-22)
The fourth inert-seam defect: v0.70.0's ghost-delete path was fully implemented and fully
unreachable. handleCustomerDelete/handleCustomerDeletePreview accepted cfg == nil, the
ghost dialog branch existed in the template JS (customer_unified.html if (d.has_config === false)) — but the Danger-zone card containing the Delete button sat inside the {{if .HasConfig}}
block (old ~L736) that also wraps the RESET card. A ghost rendered no Danger zone → no button →
dead UI. Handler tests passed because they POST directly; nothing asserted the rendered page.
Observed live on demo-vm-felhom (2026-07-22): Edit tab showed Controller Update + Geo cards only.
- Handler (
configs.gohandleCustomerUnified): newDeletablepage flag — the exact negation of the delete preview's 404 predicate (customer_delete.go:cfg == nil && no hosts && residue empty), one truth, not a lookalike. Hosts were already fetched for the Host tab; only the residue count is an extra read, and it runs solely on the ghost shape. A lookup error logs and leavesDeletable=false— fail toward hiding a destructive control. - Template: the old gate split. RESET card stays
{{if .HasConfig}}(identity-preserving re-onboarding — a ghost has no identity to preserve). Danger zone gates on{{if .Deletable}}; inside it the Block/Unblock forms gain their own{{if .HasConfig}}(blocking gates dashboard visibility of a configured customer — meaningless for a ghost). ThecustomerDeleteOpen/Submitscript moved out with the card (it was inside the old gate). Ghost shape gets a one-sentence intro prefix; the dialog already explains the rest. - Render tests (
customer_ghost_delete_render_test.go), asserting the delete form'sactionand thecustomerDeleteOpen(call site: ghost-with-residue renders Delete only (no RESET, no Block); configured customer keeps every affordance byte-for-byte (incl. the blocked→Unblock branch); nothing-left 404s (there is no renderableDeletable=falsestate —customer != nilimplies a report row implies residue > 0). Two red-proofs run and recorded inREPORT.md. - The generalized seam-wiring rule now covers template gates: any conditional UI affordance ships with a render test per branch — handler tests that POST directly prove nothing about reachability.
v0.70.0 — a deleted customer actually disappears (the ghost + its alerts) (2026-07-21)
Found while validating v0.69.0 against the live hub, on the operator's report that demo-vm-felhom
"was deleted but is still here". The delete HAD worked — config row gone, both hosts deleted, escrow
tables empty, the RESET journal complete. The customer was still on the list because GetCustomers()
builds the Customers list purely from the REPORT stream (SELECT ... FROM reports GROUP BY customer_id), and no lifecycle tier — host delete, RESET or DELETE — has ever deleted a report.
Not cosmetic: the staleness and offsite checkers iterate that same report-derived list, so the
hub kept raising offsite_stale and kept emailing the operator about a customer that no longer
exists — 10 events for demo-vm-felhom, the last one 3 days after its deletion.
Leg 3: residue (new)
The cascade is now hosts → RESET → residue → purge. The residue leg deletes, in one transaction:
reports, app_telemetry, app_log_tails, log_tail_requests, customer_notifications — plus two
rows that are not telemetry at all but credential-bearing, and outliving their customer is a
security defect rather than noise:
appliance_registrations— atoken_hash+status='delivered'row binding a box to the customer id. Deleting it returns a still-living box to the unclaimed pool on its next registration, which is the correct state for a decommissioned appliance.selfbind_tokens— an unconsumed 7-day bind token is a working path to bind a box to a customer that does not exist.
It runs BEFORE the record purge on purpose: customer_configs is the identifying descriptor and goes
last. events, notification_log, host_deletions and customer_resets still SURVIVE — the audit
trail outlives every lifecycle tier, and that rule is not relaxed here. The counter and the purge walk
one shared table list, so a table can never be counted-but-not-purged.
Ghost customers are deletable
handleCustomerDelete / handleCustomerDeletePreview used to 404 whenever the config row was
missing — so a customer deleted by any pre-v0.70.0 path could not be cleaned up by ANY operator
surface. 404 now means "there is nothing here" (no config, no host, no residue), not "there is no
config row". With no config row the offsite descriptor is unknowable, so commitCustomerReset skips
the Hetzner and descriptor legs and records skipped_no_config in the journal — never a bare
skipped, which would read as "there was nothing to do". PBS is customer-id-keyed and idempotent, so
it still runs. The dialog labels the case explicitly as a ghost and names the row count.
Tests
TestDeleteCascade_PurgesResidueAndUnlistsCustomer (residue zeroed, customer gone from
GetCustomers(), appliance + self-bind rows gone BY NAME, audit/provenance intact, journal legs
residue=ok customer_delete=ok), TestDeleteCascade_GhostCustomerIsDeletable (the exact
demo-vm-felhom shape: preview 200 with has_config:false, cascade completes, journal records
skipped_no_config), TestDeleteCascade_404WhenNothingRemains. Two more red-proofs: dropping the
residue leg leaves 5 residue rows and the customer still listed; restoring the cfg == nil 404 makes
the ghost preview 404 again. Full suite green.
v0.69.0 — customer DELETE becomes the guided full-teardown cascade (R-25b) (2026-07-21)
Implements the operator ruling of 2026-07-21. The customer page carried two half-truths: RESET
was the real teardown but refused while any host row existed, and DELETE quietly removed only
the customer_configs row (plus escrow custody) — leaving the Hetzner Storage-Box repo, the PBS
namespace + credentials, the tunnel/zone plumbing and the host rows behind. DELETE now does what its
name promises.
The cascade
POST /configs/{id}/delete (same route, new behaviour) runs three legs in a fixed order:
- hosts — every host row deleted through host-delete's own rules: an ONLINE host refuses the whole cascade (checked for every host up front, so it never half-runs) and escrow is DEMOTED to retained custody, never destroyed.
- reset — the committed RESET sequence verbatim (Hetzner FIRST → PBS → claim → descriptor → DB
purge), reached through the newly extracted
commitCustomerReset. - purge —
DeleteCustomerConfig: the customer record and all escrow ciphertext.
Nothing here is newly destructive: the cascade only sequences three operations that already existed, each keeping its own safety rules. Two invariants are load-bearing and asserted, not merely commented:
- Ruling 3 is preserved BY CONSTRUCTION — leg 2 can only run after leg 1, so the RESET sequence never sees a host row. The standalone RESET handler's 409 gate is untouched.
- Custody is purged EXACTLY ONCE, in leg 3. Leg 1 demotes; leg 2 is called with
purgeEscrow=falsesoPurgeCustomerResetDBStateleaves retained blobs alone; leg 3 is the one true purge point (v0.60.1). Both are proven from inside leg 2 by a fake that observes store state at the moment the PBS deprovision fires.
Gates (all before any write — a refused delete has ZERO side effects)
Three separate acknowledgements (ack_hosts, ack_reset, ack_purge, each must be exactly 1),
the typed customer-id, a stale-preview check (the acknowledged host count must still match
live — otherwise 409 "re-open the dialog"), and the ONLINE-host refusal. No force flag, no skip flag,
no partial-run downgrade anywhere in this path.
Resume
A failed leg retains the journal row (customer_resets, per-leg status) and the HTTP error names
the leg. Re-opening the dialog renders the incomplete journal and offers Resume; a re-run is
idempotent (leg 1 is a no-op once the hosts are gone). The acknowledgements are not cached across
attempts — a resume passes every gate again.
UI
Danger zone → Delete customer… opens a guided dialog: live inventory panel (hosts by name + status, offsite repository identifier, PBS namespace, custody state), the three consequence checkboxes, the typed customer-id field, one submit. Mid-cascade failures render the journal state. The client-side checks are convenience — every gate is enforced server-side.
Refactor (standalone RESET behaviour unchanged)
handleCustomerReset's committed half became commitCustomerReset(ctx, cfg, resetID, purgeEscrow),
returning a resetLegError (leg name + status + the exact operator-facing message). The standalone
path is byte-identical to v0.68.1: same order, same leg names, same messages, same status codes; its
suite is untouched and green. The shallow handleConfigDelete is gone — do not reintroduce a
shallow delete path.
Tests
New internal/web/customer_delete_test.go: happy-path leg ORDER (observed from inside leg 2), nine
fail-closed gate cases each asserting zero mutations and zero external calls and no journal row,
resume-after-external-failure (custody + customer survive the failure, then converge), resume is
still gated, the purgeEscrow flag's custody semantics, and a preview test asserting the inventory
names real things and leaks no secret. Five red-proofs run (ack gate, stale-preview gate,
ONLINE-host gate, leg order inverted, purgeEscrow=true) — all failed red with the wrong value
visible, then restored. Full suite green.
v0.68.1 — fix the Configuration page layout broken by the wrapper-sha field (2026-07-21)
The v0.68.0 wrapper-sha256 row wrapped itself in a <div>. The artifacts <form> IS the CSS
grid (display: grid, no inner container), so the stray </div> closed the surrounding CARD from
inside the form, and the newly-opened <div> was never closed — it swallowed the submit button and
ran to </form>. The row rendered outside the card at full page width and "Save artifact manifest"
landed inline. Reported by the operator on first use.
The field still submitted (it remained inside the form), so this was layout damage rather than data loss — but the unbalanced markup put every section below it inside the wrong container.
Fixed by making the row plain grid cells (grid-column: 2 / 4 for the input and the hint), with no
nested elements at all.
There was no render assertion on this form, which is why a hand-edit could break it silently.
TestConfigurationArtifactsForm_Structure now asserts the field is inside the form, the submit
button has not escaped, the form contains zero <div>s, whole-page div balance holds, and the
sections that render after it still exist. Red-proofed by restoring the broken shape.
v0.68.0 — a credential re-issue finally re-arms the box (R-39 fleet fix) + wrapper drift is visible (R-50b(a)) (2026-07-21)
Coupling, stated honestly: this release is SAFE for agents at 0.90.0 — the new descriptor field is an unknown JSON key to them; they drop it and behave exactly as today (inert, not breaking). The re-arm and auth-honesty guarantees require agent >= 0.91.0. Raise MinAgent to 0.91.0 only after the fleet's agents have self-updated.
The defect (R-39, fleet half)
An ep0 credential re-issue re-keys the secret of an existing token. token_id, fingerprint,
datastore and namespace all come back byte-identical — only the side-table host_pbs_secrets
row rotates. The agent's re-apply trigger is a change in the descriptor content hash
(felhom-agent internal/pbsdr/manager.go descriptorHash). Same hash → the converged agent
short-circuits → the fresh secret is never consumed → the box keeps presenting a revoked credential
→ 401 forever, while both tiers report applied. Proven on the N100 2026-07-18: the agent's
consumed-failed.json hash was byte-identical to the marker.json written two minutes before the
re-issue.
The fix
host_pbs_secrets.generation— a monotonic per-host counter advanced by every fresh mint and by nothing else, stamped into the descriptor assecret_generation. That is now the only field a re-key moves, and it is what re-arms the agent.- A re-stage deliberately does not advance it: it re-arms the same secret, the descriptor content genuinely has not changed, and a bump would cause a pointless agent refetch loop.
omitemptyis load-bearing — emitting a zero into every pre-existing descriptor would itself be a fleet-wide spurious re-apply.- Deviation from the spec, deliberate: the brief said to reuse "the new row's id … no schema
change". There is no row id — the table is keyed by
host_idand UPSERTed last-write-wins, so a new row never exists, andcreated_atcollides for two mints in one second. An additive counter column (existing idempotentALTER TABLEidiom) is the only monotonic source available.
pbsdrhealgains anauth_failedtrigger — a NEW trigger in the existing machine, not a new machine. A box whose credential PBS rejects is escalated to a fresh mint (never a re-stage: that re-feeds the secret PBS just rejected), through the existing damping — a 401 flap must not become a secret-minting chain. This closes the loop end to end: agent proves the 401 → hub re-keys → generation advances → descriptor hash moves → agent re-consumes.consumed_athonesty gauge — a staged secret left unconsumed past a 15-minute grace while the box reportsappliedis surfaced loudly with its own event. This is the exact July-18 fingerprint and a disagreement no single tier can detect alone. Deliberately a surface, not a heal: auto-re-issuing here would mint a second secret on top of an unconsumed one — the mint/consume race R-39(a) already recorded.- Corrected a comment that stated a falsehood:
ReissuePBSDRclaimed it refreshed the descriptor "with the NEW token_id/fingerprint". That is false for a re-key, and believing it is why nobody expected the descriptor to come back identical.
R-50b(a) — wrapper drift is answerable
ArtifactManifest.WrapperSHA256 + an operator field. Unlike the agent binary and the golden, the
PBS-DR wrapper is installed from raw/branch/main — unversioned, unpinned, absent from every
manifest — yet it is root-owned 0755 and the pinned sudoers vector. Agents (>= 0.91.0) report the
installed file's hash and the host page surfaces a mismatch. An unknown on either side reads as
quiet, never as drift — lighting every host amber on rollout day is how a warning becomes noise.
This does not fix the delivery channel; that stays R-50b(b)/(c).
Tests
Store-level generation monotonicity, per-host isolation and restage-leaves-it-alone; descriptor
byte-change, omitempty and sibling-key round-trip; a flow-level test driving ReissuePBSDR
against a fake that models a real re-key; auth_failed escalate/debounce/recovery; the honesty
gauge incl. its grace window, the restage edge, the consumed case and the honest-stuck case; wrapper
drift incl. both unknown directions. Two red-proofs run at the assertion level (not the
compiler): removing the generation stamp makes the flow test fail with both byte-identical blocks
printed; removing the auth_failed arm makes the escalation tests fail with reissues=0.
v0.67.0 — the hub stops keeping things to itself: auto-minted self-bind link, post-RESET staleness, unprovisioned-offsite warning (2026-07-18)
Four small items, each one a case where the hub already knew something and said nothing. Green:
go build ./... && go vet ./... && go test ./... all pass.
-
(a) The self-bind link is minted automatically — at customer creation AND at RESET completion (R-36 sub-item). The box's console banner tells the customer to open „az e-mailben kapott link"; until now that email existed only once the operator remembered to press Send self-bind link, so the banner could be instructing someone to look for something that did not exist. During the 2026-07-18 rehearsal the box sat in pairing mode for ~11.7 minutes waiting on exactly that.
handleSelfBindLinkSend's body was extracted into a sharedmintAndSendSelfBindLinkcore so the button and the two auto-mint call sites cannot drift apart on the honesty rules: F1 (no registered address → mint nothing, because a link nobody can receive is worse than none) and F2 (send failed → delete the token, never leave it silently live). The auto-mint wrapper never fails the operation it rides on — a customer create that provisioned Cloudflare, offsite and PBS must not 500 because a courtesy email bounced; every outcome is logged instead, and the operator can still re-send from the Setup tab. Gap found and closed while wiring this:PurgeCustomerResetDBStatedoes not clearselfbind_tokens, so a capability link minted before a RESET would have stayed live across it. A successful mint already replaces it (minting is delete-then-insert, single-active), but the skip paths would not have — so the wrapper now clears stale tokens on those paths too. The invariant is now: after auto-mint runs, the only live link is one it just issued, or none. -
(b) Post-RESET staleness banner (R-37). When a RESET completed after the newest report, every health figure on the customer page describes a lifecycle that no longer exists — and the page went on showing pre-RESET warnings as if they were current. It now says so, quoting the customer-facing phrasing („RESET óta nincs adat") and the reset's timestamp. Deliberately narrow: an in-flight reset does not trigger it (only a completed one), and it clears itself the moment a report arrives. Ties resolve to stale — SQLite timestamps are second-resolution, and a same-second report almost certainly arrived just before the reset destroyed what it describes; erring the other way would hide the banner exactly when it matters most.
-
(c) Unprovisioned-offsite warning (R-36 interim).
offsite.enabled == truewith no descriptor (type == "") is a real, stable, silent state: provisioning is Save-triggered (applyOffsite), and the re-enroll auto-re-issue deliberately skips an unprovisioned target, so nothing self-heals it. The page now names the state and the fix (press Save once, then verify), reusing the exact predicate the offsite re-issue handler already refuses on. -
(d)
pbsdr_reissuedrendered an EMPTY flash box. The flash key had no branch in the template, so re-issuing PBS credentials showed the operator a success box containing nothing — observed live on 2026-07-18. It now describes what was staged and carries the R-39 caveat: confirmpvesm statusshows the entry active, because a converged agent can reportappliedwhile the storage still authenticates 401. -
Styling: new
.flash-warn(amber,--warn/--warn-dimtokens) for the deviation tier between success and error — per the exception-color principle it appears ONLY on deviation, never on a healthy page. -
Tests (
customer_state_banners_test.go,selfbind_automint_test.go) assert each banner is absent in the nominal cases as well as present in the deviating one — a banner that renders unconditionally is worse than none, because operators stop reading it. Auto-mint covers F1, F2, the no-mailer case, and the pre-RESET-token invariant. Both red-proofed: deleting thepbsdr_reissuedbranch reproduces the original empty-box bug, and neutering the staleness predicate fails the banner assertion. New store accessorCountSelfBindTokens(read-only, keyed by a customer id the operator already knows) makes the single-active invariant assertable.
Not in this train: the R-39 hub-side generation-bump fix the pre-travel task made conditional.
Its condition was refuted — SetHostDesired bumps unconditionally and applyPBSDR is
idempotent as documented; the real mechanism is the agent's descriptor-hash convergence, which needs
its own spec. Nothing was improvised here.
v0.66.0 — Customer self-bind (R-27 slice 1): tokenized capability link + public two-factor /bind/ page (2026-07-17)
Lets a customer bind their OWN freshly-installed appliance without the operator. Until now every
box booted from the universal secret-free ISO (R-21 slice C) had to be bound by the operator on the
Hosts page; this adds the self-service path. Viktor's three rulings, each honoured verbatim:
(a) "only their own visible" → the customer proves possession with the console pairing code
shown on the box screen — no appliance list is ever rendered on any public surface; (b) first-box
entry → an operator-sent 7-day tokenized capability link over Hungarian email (the claim-engine
delivery pattern, a sibling sender — NOT routed through the claim engine); (c) lockout after 5
failed attempts → the token locks and the page says "call support". Wrong code and wrong
passphrase produce one identical generic failure (no oracle); an expired link falls back to
operator-bind, unchanged. Does not touch the controller or agent. Green: go build/vet/test
(full hub suite + 9 new self-bind tests), all 4 red-proofs verified red-then-green, hub confirm gate.
- Pairing code (Part 1).
POST /api/v1/appliance/registernow returns an additivepairing_code(6 chars, ambiguity-free alphabet,ABC-234display) minted once at first registration and stable across the idempotent re-register/upsert (backfilled if a pre-existing row had none). Persisted onappliance_registrations.pairing_code; shown in the operator Hosts "Unclaimed appliances" table. The bootstrap (felhom-bootstrap.sh, ISO v1.20.0) parses it and prints a Hungarian console banner to/dev/consoleso the customer can read it off the screen. Old ISOs ignore the field; an old hub omits it and the banner prints nothing — additive both ways. - Capability token (Part 2). New
selfbind_tokenstable:sha256(token)at rest (never the token), single-active per customer (a re-mint deletes the prior row in one tx — an old link dies the instant a new one is sent),attemptslocking at 5, one-shotconsumed_at,emailed_athonesty. Operator "Send self-bind link" button on the customer Setup tab (POST /customers/{id}/selfbind-link) mints a 256-bit token and emailshttps://hub.felhom.eu/bind/<token>(Hungarian, adult tone, no emoji, names both factors + the 7-day + 5-attempt limits). F1 (no registered email → nothing minted, LOUD flash) and F2 (email send fails → the just-minted token is deleted, not left silently live) are both honest. - Public bind page (Part 3) — THE TRAP (§9.2). One new public prefix
/bind/, exempted from operator auth AND CSRF at the two gate sites the/loginexemption occupies, via a single predicateisPublicBindPath(matched tightly: trailing slash → no sibling like/bindsecret; the ServeMux..-cleans before we see the path → no traversal reach; the handler also rejects a token containing/).GETrenders form/consumed/locked/expired;POSTnormalizes both inputs, compares both factors unconditionally (constant-time passphrase vs the customer's retrieval passphrase; the ONE bindable appliance carrying the console code), then decides — identical generic failure either way. On success: the sameBindAppliancethe operator uses, a provenance event with sourcecustomer_selfbind, one-shot consume, and "A doboz kb. egy percen belül folytatja a telepítést." The box's ~30 s appliance poll picks up the delivery. Own per-IP rate limiter; the page is self-contained (it cannot link/style.css, which is itself operator-gated). No hub customer-login/session was built — the URL capability token IS the auth model; a cross-site POST without both secrets only burns attempts (accepted + documented). - Tests + red-proofs (Part 4). Scenarios A–F + F1/F2 (9 tests). The 4 red-proofs were each
applied and confirmed to turn exactly their scenario red, then reverted green: lockout removed → C1;
oracle introduced → B;
/bind/prefix widened (drop the slash) → E (and D); single-active DELETE dropped → C4. The passphrase is never logged/echoed/persisted; only attempt COUNTS are logged (self-bind attempt N/5 … token <8hex>…); the raw link token never enters logs or events. - GC verdict (spec §3): there is no appliance-staleness GC in the hub (
applianceStaleAfteris a DISPLAY badge only;pruneAll/PurgeExpiredLogBundlestouch reports/log-bundles, not appliances or selfbind tokens). The 7-day token TTL therefore stands alone and needs no reaper — single-active-per-customer means at most one row per customer, superseded rows are deleted on re-mint, and an expired row simply reads as expired (no security or storage pressure).
v0.65.0 — PBS DR storage visibility (ep0 usage op) + Offsite tab split (Restic / PBS DR) + dual dashboard gauges (R-5) (2026-07-17)
Makes the PBS DR storage visible like the restic pool box already is (v0.64.0), the two clearly
differentiated. The scoping correction: "the restic box" and "the PBS box" are NOT two Hetzner Storage
Boxes — restic = subaccounts on the shared Hetzner Storage Box (Hetzner API, v0.64.0); PBS DR = a
PBS datastore (felhom-offsite) on the ep0 endpoint VM (NO Hetzner API). This adds the hub's read of the
PBS datastore fill via Option A (Viktor-ruled): a new read-only usage op on the felhom-tenantsync
ep0 forced command (the structural twin of the existing fingerprint op), polled by a new hub checker on
the same 15-min throttle; splits the Offsite page into Restic / PBS DR tabs; and puts two dashboard
gauges (restic %, PBS %). READ-ONLY against ep0 and Hetzner. Green: go build/vet/test + bash -n +
the script harness; hub confirm gate OK.
- Phase-0 probe (gate, PASSED): on ep0 (PBS 4.2.3),
df -B1 --output=size,used,avail <datastore path>(path fromproxmox-backup-manager datastore list --output-format json) yields the datastore total/used/avail in bytes (live: 39990112256 / 7627939840 / 30686326784 → ~19%), read-only, in the existing sudo context, no admin token. (PBS 4.2 has no nativedatastore usagecommand.) scripts/felhom-tenantsync.sh→ v1.2.0: a read-onlyusageshort-circuit (before the admin-token generation, likefingerprint) →{"status":"ok","total","used","avail"}. No customer_id, no admin token, NO mutation.tenantsync.Client.Usage()(internal/tenantsync/client.go):BoxUsage{Total,Used,Avail}+ the op; an endpoint ≤ v1.1.0 answersbad_request "unknown op"→ typedErrUsageUnsupported(the graceful-degradation signal).monitor.PBSDRBoxChecker(internal/monitor/pbsdr_box.go, new): clones OffsiteBoxChecker over ausageReaderseam (the tenantsync client) — 15-min throttle, cachedPBSBoxSnapshot, escalation-onlypbsdr_box_fillon the customer-less"pbsdr-box"scope (operator channel only, no SaveEvent), recovery re-arm. FILL ONLY (PBS uses namespaces, not quotas — no oversubscription). THREE states:ok(bands drive),unavailable(ErrUsageUnsupported — expected pre-update, neutral, NO alert, logged once, the gauge shows n/a),degraded(exec failed — keep last snapshot, no band transition).- Config + wiring:
Alerting.PBSDRBoxFill{Warn,Crit}Percent(default 80/90, independently tunable); the checker is built ONLY when the tenantsync client exists (shares it), registered in the 60 s sweep, snapshot handed to the web server. Graceful degradation: the hub deploy is INDEPENDENT of the ep0 update — a hub v0.65.0 against an ep0 still on v1.1.0 shows the honest "n/a", lighting up on the next poll once ep0 is updated (no hub redeploy). - Web (
internal/web/pbsdr_box.gonew,offsite.go,templates/offsite.html,dashboard.html,style.css): the Offsite page splits into Restic (the v0.64.0 pool-box panel + per-customer rows) and PBS DR (a new datastore panel — capacity/used/fill bar with band + the endpoint cards, which belong here: the endpoint IS the PBS host) hash tabs (server-rendered, no JS dependency for the data). The single dashboard tile becomes two gauges — RESTIC (pct·ratio) and PBS DR (pct; "n/a" when unavailable) — each band-colored, each linking to its tab. - Tests + red-proofs: 10 Go tests (Usage parse + unknown-op→typed; checker throttle/bands/pbsdr-box
operator-only/unavailable-no-alert/degraded-keeps-last; PBS panel render × ok/unavailable/not-configured)
- a bash script harness (usage JSON + exit 0 + zero mutation + provision regression). Red-proofs (run-fail-restore): the op emitting a mutation → the harness zero-mutation assertion fails; the escalation-only guard removed → in-band re-emit fails; unavailable driving a band → the no-alert test fails. All confirmed red, then restored.
v0.64.0 — offsite pool-box aggregate: fill, oversubscription, per-customer bars, operator alert (R-5) (2026-07-17)
Ships R-5: the operator sees the shared pool box's real state on the hub — total box fill vs
capacity, Σ(shared soft quotas) vs capacity (the oversubscription ratio), per-customer usage/quota
bars, and a box-level operator alert (fill % + oversubscription ratio) riding the existing
dispatcher's operator channel. Per-customer fill alerts already existed (OffsiteChecker, 90/95% of each
quota); the box-level aggregate was the gap — the operator's early warning that the pool itself is
filling, before any single customer breaches. All READ-ONLY against Hetzner (GET only). Green:
go build/vet/test all pass; hub confirm gate OK.
- Phase-0 probe (gate, PASSED): one authenticated GET of the live pool box (611714) pinned the API
shape — capacity is
storage_box_type.size(1 TiB for bx11), usage is astatsobject (size/size_data/size_snapshots), all bytes; our token reads it (200). The type extension mirrors it. hetznerapi(hetznerapi.go,fake.go): additiveStorageBoxType+StorageBoxStatssub-structs onStorageBox(no existing field/method changed); fake carries them + aGetBoxCallscounter. Golden decode test against the (redacted) probe capture.monitor.OffsiteBoxChecker(offsite_box.go, new): the OffsiteChecker-sibling for the box as a whole — fetch-throttled (one Hetzner GET per 15 min; ≈4/hour, never per-sweep or per-page-load), cachedBoxSnapshot, escalation-only emits with silent recovery re-arm. Two independent signals: FILL (used/capacity, warn 80% / crit 90%) and OVERSUBSCRIPTION (Σ shared+enabled quotas / capacity, warn 2.0×) — both can fire, neither masks the other. Σ(quota) is read from the authoritative ConfigJSONDescriptor(offsite.ReadDescriptor, new), NEVER the report echo; dedicated + disabled customers excluded. Events carry the customer-less scope"pool-box"→ operator channel ONLY (processCustomerno-ops on it) and are NOT SaveEvent'd (no customer row to key them). A failed fetch keeps the last snapshot marked degraded — missing data never becomes 0% and never drives a band transition.- Config (
cmd/hub/main.go):Alerting.OffsiteBoxFillWarnPercent(80) /OffsiteBoxFillCritPercent(90) /OffsiteOversubWarnRatio(2.0), plumbed likeStorageFill*. The checker is constructed inside the existingHETZNER_TOKENbranch (shares the client), registered in the 60 s sweep, and its snapshot handed to the web server. Thresholds are Claude's encoding of the starter suggestion — Viktor's ruling pending; the named keys are the one-line flip. - Web surfaces (
web/offsite_box.gonew,templates/offsite.html,dashboard.html,style.css): an Offsite-tab panel (capacity, used with data/snapshot split, fill bar, Σ quotas + ratio, fetched-at, + per-customer rows sorted by usage — shared with a usage/quota bar, dedicated listed without one, no-report customers show "no usage reported yet") and a compact Dashboard tile (fill% · ratio, band-colored, linking to /offsite). The web layer reads only the cached snapshot — it NEVER fetches. Nil provider → both render an honest "not configured". Exception color (neutral/amber/red, no green). - Tests + red-proofs (run-fail-restore): 10 new tests (throttle ≈4-not-≈60, Σ/ratio truth table, band transitions incl. in-band no-re-emit + recovery re-arm + oversub independence + pool-box operator-only scope, failed-fetch honesty, zero-capacity guard, golden decode, panel render × "not configured"/"with data"/"no usage reported yet"). Red-proofs: (i) drop the throttle → ≈60 calls; (ii) sum dedicated/disabled → wrong Σ; (iii) drop the escalation-only guard → in-band re-emit; (iv) zero the snapshot on a failed fetch → lost last-known. All confirmed red, then restored.
v0.63.0 — system-initiated immediacy: wire the proven poke/bump notifiers into every mutation site that lacked one (2026-07-17)
The immediate-sync arc (Dir-1 trigger, Dir-2b wait channel, Dir-2a agent poke) covered only
operator-initiated desired-state changes. System-initiated mutations still bumped the
generation silently, so a freshly onboarded box waited a full agent tick (≤15 min) for state the hub
had already minted — observed live during slice-C onboarding. This wires the existing, live-proven
notifiers (poke.Notifier for the agent plane, intent.Hub.Bump for the controller plane) into
every system-initiated site that lacked one. No new mechanism — call-site wiring only.
- Agent-plane pokes (
internal/web/pbsdr.go):PBSDRAutoProvision(the exact observed lag — the WG-registration hands-free provision) now pokes on success;ReissuePBSDR(the shared core, which also gives the pbsdrheal reconciler's escalation its immediacy with zero reconciler changes) and the operator buttonhandlePBSDRReissuepoke after their descriptor bump. All fire ONLY after the successfulSetHostDesired, never on a blocked/error path. - Agent-plane Poker seam (
internal/api/handler.go,internal/api/wg.go): a new nil-safePokerinterface (PokeHost/PokeAllHosts, satisfied by*poke.Notifier) +SetPoker.handleAdminSetDesiredStatepokes the target host after a successful admin desired-state write;handleAdminSetOperatorPeerfires a fleetPokeAllHosts— but only whenBumpAllHostGenerationssucceeded (fire-after-commit). - Controller-plane bump (
internal/api/handler.go):reissueOnReenroll(the clean-slate F2 claim + F3 offsite re-issue) nowintentHub.Bumps the customer so a long-polling controller wakes in seconds instead of on the 15-min cycle. Nil-guarded; one unconditional bump (coalesced, over-bump harmless). - main.go: one
poke.Notifierinstance now feeds BOTH planes' system sites —webServer.SetPoke(n)andapiHandler.SetPoker(n); startup log: "web + api admin seams armed". - Deliberate non-sites (unchanged, documented in REPORT audit table): WG peer register (box tunnel
doesn't exist pre-fetch — a poke is undeliverable; the agent fast-tick SECONDARY owns this leg), WG peer
delete (the mutation removes the transport), and the pbsdrheal Restage path (no generation bump →
the agent's 60 s pbsdr ticker is the pickup path; a poke there is a verified no-op).
internal/pbsdrheal/is byte-unchanged. - Tests + red-proofs (run-fail-revert): 10 new non-hollow tests. Web (async, channel-synchronized
fake sender): auto-provision pokes the resolved WG /32, blocked-precondition pokes nothing, reissue-core
- operator-button poke, a failed reissue pokes nothing. API (synchronous fake Poker): admin-set pokes
the target host only (0 on invalid-JSON), operator-peer fires one fleet poke, nil-seam no-panic; re-enroll
advances the intent generation, nil-hub no-panic. Red-proofs demonstrated one representative removal per
group (A/B web pokes, C admin poke, D re-enroll bump) — each FAILED red, then restored. Green:
go build/vet/testall pass.
- operator-button poke, a failed reissue pokes nothing. API (synchronous fake Poker): admin-set pokes
the target host only (0 on invalid-JSON), operator-peer fires one fleet poke, nil-seam no-panic; re-enroll
advances the intent generation, nil-hub no-panic. Red-proofs demonstrated one representative removal per
group (A/B web pokes, C admin poke, D re-enroll bump) — each FAILED red, then restored. Green:
v0.62.0 — R-21 slice C: the universal ISO — unclaimed-appliance registration + operator bind + one-shot delivery (2026-07-17)
The hub half of the universal, secret-free bare-metal ISO. A box booted from the generic ISO registers itself as an UNCLAIMED APPLIANCE; the operator binds it to a customer on the Hosts page; the hub delivers the customer-id + retrieval passphrase on the box's next poll, ONCE. The distributed ISO carries no customer secret (§4.4).
- Store (
internal/store/appliance.go, new):appliance_registrationskeyed by (uuid, mac_set) — serials are unusable (N100 DMI "Default string") and cheap boards duplicate SMBIOS UUIDs, so the MAC set is the tiebreaker (same uuid + different mac-set = distinct appliance).token_hash= sha256 of the appliance token (the token itself is never stored).RegisterAppliance(idempotent upsert; sticky-discard),ApplianceByToken,BindAppliance,MarkApplianceDelivered(atomic one-shot bound→delivered),DiscardAppliance(invalidates the token),ListUnclaimedAppliances. The table's own timestamps ARE the pre-bind provenance (no customer to scope an events row to yet). - API (
internal/api/appliance.go, new):POST /api/v1/appliance/register— the ONE unauthenticated endpoint, per-IP rate-limited, returns a random 256-bit appliance token.GET /api/v1/appliance/poll(Bearer token): unknown/discarded → 404 (no oracle), unbound → 204, bound → 200 + credentials (consumed once), delivered → 410. The passphrase is read live fromcustomer_configs(plaintext, as the day-0 command already needs it) and never logged. - Web (
internal/web/appliances.go, new): the Hosts page grows an "Unclaimed appliances" section (uuid, MACs, hw, SSH host-key fingerprints, first/last seen, stale >7d badge) with BIND (customer picker showing host counts — display only, never a gate) and DISCARD. Bind stages the delivery + emitsappliance_bound; delivery emitsappliance_credential_delivered. - Red-proofs (run-fail-revert): the one-shot delivery (defeat the bound→delivered flip → second
poll re-delivers the passphrase → FAIL) and register idempotency (drop the upsert → duplicate/UNIQUE
violation → FAIL), both proven red then restored; plus 404-no-oracle + sticky-discard, bind
staging/refusal, and the render test. Green:
go build/vet/test; hub confirm gate OK.
v0.61.0 — Customer RESET: the middle lifecycle tier (2026-07-17)
One operator action returns a customer to pre-first-install: every OPERATIONAL trace dies (offsite
repo, PBS namespace + backups, DR recipe, one-time secret, claim state, retained escrow custody), while
identity and the basic config survive (the customer_configs row, all provenance rows, and the
audit-event stream). It sits between the two existing tiers — host delete (< RESET) and customer
Delete (> RESET, the one true purge point). Viktor's rulings: (1) destroying retained escrow custody
gets its own separate acknowledgment; (2) RESET clears claim state (a fresh code next onboarding);
(3) RESET refuses while any host row exists (delete hosts first — reset never deletes hosts); (4)
the confirm surface shows a live-counted inventory.
Orchestration discipline (spec §3): external teardown FIRST, DB purge LAST (publish-last), every leg
idempotent → a partial run is simply re-run from the top; a failed external leg is a clean journal
entry and the DB purge (which erases the descriptors that say what still needs tearing down) is
withheld until every external leg is ok. Provenance + events are NEVER wiped.
- Store (
internal/store/customer_reset.go, new):customer_resetsjournal table (per-attempt, per-leg status, resumable);CustomerResetInventory(live counts: hosts, retained blobs via the F-14host_deletionsUNION, dr_recipe/one-time-secret/claim presence);Start/UpdateResetLeg/Finish/ LatestCustomerReset;PurgeCustomerResetDBState(ack-gated escrow-blob delete + one-time-secret, dr_recipe, log bundles — never touches identity/provenance/events);DeleteClaimprimitive. - Claim (
internal/claim/engine.go):ResetToUnclaimedDELETES the claim row soEnsureIssuedmints a fresh first code on the next onboarding (no parallel revoked-flag, no stale generation). - Offsite (
internal/offsite/offsite.go):DeprovisionDELETES the labelled sub-account/box (idempotent — label-lookup,len==0= already gone);OffsiteIdentifier(preview name);ClearProvisionedDescriptor(keeps the tier CHOICEenabled/type/quota/box_type, drops every provisioned field). PBS (internal/tenantsync/client.go+scripts/felhom-tenantsync.shdeprovisionop): destroys the customer's namespace + all backup groups + token; the sharedfelhom@pbsuser is never touched; idempotent. - Web (
internal/web/customer_reset.go, new):GET /configs/{id}/reset→ live inventory JSON;POST→ the orchestration (precondition + typed-id + escrow-ack gates BEFORE any write/external call). A distinct amber RESET card on the customer page (separate from the red Danger-zone Delete), with the typed-id confirm + the separate escrow-custody ack row. - Red-proofs: ack-gate (defeat → reset proceeds & destroys blobs → FAIL); partial-failure
resumability (purge-not-withheld → DB purged despite external failure → FAIL); both proven red then
restored. Plus store ack-gating, journal round-trip, offsite Deprovision idempotency + descriptor
clear, and the RESET-card render test. Green:
go build ./... && go vet ./... && go test ./....
v0.60.1 — host deletion DEMOTES escrow custody (never destroys) + S6b obsolete (2026-07-17)
Closes the deletion-path gap in v0.60.0's review: DeleteHost(deleteEscrow=true) was still DELETING
escrow rows (the same-customer reinstall flow funnels the operator straight into that tick).
Principle (Viktor's standing ruling): host deletion is a lifecycle event — blob custody survives it;
the customer Danger-zone Delete is the one true purge point. Green:
go build ./... && go vet ./... && go test ./....
- Scenario A — host delete demotes, never destroys.
DeleteHost(deleteEscrow=true)now DEMOTES the currenthost_escrowrow intohost_escrow_superseded(copy-BEFORE-delete, same tx) and SPARES existing superseded rows — no operator path through host lifecycle can lose a blob. Reuses THE one escrow row-copy routine (demoteCurrentEscrowTx, also used bySaveHostEscrow). The F-14 provenance row + gate semantics are unchanged (wording updated: demotion, not destruction). Edge: no escrow row → unchanged; theErrHostEscrowPresentrefusal without the flag is unchanged. Red-proofTestDeleteHost_DemotesEscrowNeverDestroys. - Scenario B — customer delete is the purge point.
DeleteCustomerConfig(which previously deleted ONLY thecustomer_configsrow) now, in one tx, purgeshost_escrowANDhost_escrow_supersededfor all the customer's hosts — INCLUDING already-deleted hosts (resolved via the F-14host_deletionsprovenance) so a host-delete-then-customer-delete ordering leaves nothing orphaned. Danger-zone copy states it. Red-proofTestDeleteCustomer_PurgesEscrowCustody. - Scenario C — wording. The host-delete escrow checkbox now reads "Move key escrow to retained
custody (required when escrow present)…"; the refusal message + customer Danger-zone copy match.
Guard test
TestHostDeleteEscrowLabel_DemotionWording. - Scenario D — S6b verdict (docs): OBSOLETE. Re-enrolling an existing host_id upserts cleanly
(
UpsertHostON CONFLICT DO UPDATE,store.go;handleAdminCreateHosthas no duplicate refusal) + the v0.57.0 re-enroll arc auto-fires the re-issues → no manual stale-host deletion needed before re-enroll. Scenario A also makes the funnel harmless either way. ROADMAP R-3 refined. - Scope: hub-only; no controller/agent change; ACK assembly + upload supersede path untouched.
Deploy: bump
manifests/hub.yamltag to0.60.1and sync.
v0.60.0 — offsite continuity Part B: superseded-escrow retention (data-first) (2026-07-17)
Closes the data-loss half of the reinstall-orphaned-repo incident: SaveHostEscrow's destructive
ON CONFLICT overwrite meant a new escrow blob DESTROYED the old passphrase's only copy — so 18
snapshots keyed under the old password became unrecoverable. Viktor's ruling (data protection first):
retain superseded blobs so the old passphrase stays customer-R-recoverable. Pairs with controller
v0.142.0 (Part A orphaned-repo guard). Green: go build ./... && go vet ./... && go test ./....
host_escrow_superseded(new table) +SaveHostEscrowrewrite. On upload, if the current row seals a DIFFERENTrestic_pw_sha256, the old row is COPIED into the history table (in one tx) BEFORE the current row is overwritten; a same-sha re-upload (idempotent re-ceremony) refreshes the current row and creates NO supersede row.SaveHostEscrownow returnssuperseded bool. Retain ALL (no pruning — the blobs are tiny + R-encrypted, custody unchanged); the hub still never decrypts. ACK/restore-serving read the CURRENT row (GetHostEscrow) — unchanged. NewCountSupersededEscrow/ListSupersededEscrow(the latter seeds the future R-26 recovery flow).DeleteHost(deleteEscrow=true)also drops the retained rows.- Surfaces. Upload handler emits the hub-internal
escrow_supersededaudit event + logs the retained count; the operator host-detail DR/Backup panel shows "N superseded escrow blob(s) retained". Registeredoffbox_repo_orphaned/offbox_repo_reset(controller v0.142.0 pushes) inallowedEventTypes+customerMessages. - Red-proof
TestSaveHostEscrow_RetainsSuperseded(pre-fix destructive overwrite → old blob gone → FAIL; fixed → retained + retrievable; same-sha idempotent). - Deploy: bump
manifests/hub.yamlimage tag to0.60.0and sync.
v0.59.0 — Direction-2a: agent-plane immediate-sync poke sender + ep0 felhom-poke surface (2026-07-16)
Implements the AGENT-plane half of documentation/audits/SPIKE-immediate-sync-transport-2026-07-16.md
option (a): an operator agent-plane change (a pbsdr descriptor, a MinAgent floor) now nudges the box
in seconds via a CONTENTLESS UDP poke relayed hub → ep0 forced-command → wg0-origin → the agent's
poke listener (felhom-agent v0.89.0). The spike reserved this transport for the agent plane and
measured it at ~0.42 s/poke. Complements v0.58.0's internal/intent long-poll wait channel
(Direction-2b, the CONTROLLER plane): intent is customer-keyed and wakes the config puller; poke is
host-keyed and nudges the agent. Both are fire-and-forget; the 15-min report cycle stays the
guarantee. Pairs with felhom-agent v0.89.0 (the listener).
internal/poke— the SSH poke sender + notifier. Third structural sibling ofinternal/wgsync/internal/tenantsync: a pinned-host-key (ssh.FixedHostKey, algorithm-pinned) in-process SSH client that reuses the peersync endpoint + host key, its OWN forced-command key.Client.Poke(ctx, boxWGIP)refuses any target outside10.77.0.0/24BEFORE dialing, then SSHes to ep0'sfelhom-pokewith the box's WG /32 as the command string (→$SSH_ORIGINAL_COMMAND); the forced command sends one empty UDP datagram to<ip>:51822.Notifier.PokeHost/PokeAllHostsare fire-and-forget (detached goroutine, nil-receiver-safe) — a poke NEVER blocks or fails the operator save; a missing peer / SSH error is logged and the report cycle reconciles. Tests: resolved-IP send, no-peer/store-error no-send, nil no-op,Pokepre-dial non-WG refusal.- Wiring (
internal/web/server.go,internal/web/pbsdr.go,internal/web/configs.go,cmd/hub/main.go):Server.SetPoke;applyPBSDRfiresPokeHost(host.HostID)after each generation-bumping descriptor save (disable, storage-id/re-enable change, fresh provision);handleSetArtifacts(the MinAgent-floor / vouched-agent save) firesPokeAllHosts(). Source note (spec landmark vs source): the artifact-manifest save does NOT itself bump per-host desired generation — the agent self-update dispatches via signed-ops on the next report — so the fleet poke there accelerates the next report cycle where the floor is applied, rather than delivering a desired-state delta. The sender is env-configured (POKE_SSH_KEY_FILE, reusingWG_ENDPOINT_SSH_ADDR/_HOSTKEY/user); absent Secret → "agent-plane poke disabled". - ep0 surface (
scripts/felhom-poke.shv1.0.0 +documentation/runbooks/offsite-endpoint.md§11): a NON-root (felhom-peersync, no sudoers grant — a datagram needs no privilege) forced-command that validates$SSH_ORIGINAL_COMMANDto the WG /24 and sends one empty datagram from wg0. Contentless + confined (the WG kernel independently refuses non-peer /32s — spike P1 EKEYREJECTED). Port 51822 is a shared cross-repo constant. manifests/hub.yaml:POKE_SSH_KEY_FILEenv + optionalSecret/agent-pokemounted 0400 at/etc/hub-secrets/agent-poke/key(the private key stored out-of-band, per §11). Bump the image tag to0.59.0and sync.
v0.58.0 — Direction-2 immediate-sync: the hub→box "sync now" wait channel (2026-07-16)
Implements option (b) of documentation/audits/SPIKE-immediate-sync-transport-2026-07-16.md: an
operator action on the hub now reaches the box in seconds instead of on the next ~15-min report
cycle. The box holds a hanging authenticated GET /api/v1/wait over the existing outbound ingress;
the hub completes it the instant any operator intent lands for that customer. The box then fires its
ordinary out-of-cycle report — the ACK delivers config/escrow/claim/floor through the UNCHANGED
machinery. The 15-min cycle stays the reconciliation backbone; every wait failure degrades to it.
Pairs with controller v0.140.0 (the long-poll client). Ground truth #1 holds: the box pulls even the
wake-up; the hub never connects inbound and no state ever rides the wait response.
internal/intent— the in-memory operator-intent notifier. A per-customer generation counter with a waiter registry:Bump(customerID)advances the generation and wakes every registered waiter (a burst coalesces into ONE completion carrying the LATEST generation — a counter, not a per-bump queue);Wait(ctx, customerID, lastSeen, maxHold)returns the instant the generation differs fromlastSeen, on ctx-cancel, onmaxHold, or onClose. A pre-register gen-check closes the bump-before-connect race (a bump is never lost). In-memory BY DESIGN — a hub restart resets generations; the box compares with!=, so a restart costs exactly one harmless full-state report, never a storm. No persistence, no store schema. Red-proofs: counter-vs-queue (return the as-of-register snapshot →TestWait_CoalescesBurstToLatestGenfails) and the race-closer (drop the pre-register check →TestWait_RaceCloser_BumpBeforeWaitNotLostfails); both run-fail-reverted.GET /api/v1/wait(api/wait.go). Authed viacheckAuthCustomer; per-customer only (a global operator key → 400; the customer is resolved from the key, nocustomer_idparameter is accepted, so A can never observe B). Holds up to 240 s, writing a 25 s heartbeat newline while it waits. nginx'sproxy_read_timeoutis measured BETWEEN upstream reads, so the heartbeat keeps the default 60 s from ever firing — no ingress annotation / manifest timeout change is needed (the transport spike measured that ceiling; §13 proves the heartbeat defeats it live). The response is contentless — a single{"gen":N}line. The connection's write deadline is lifted per-request viahttp.NewResponseController().SetWriteDeadline(the globalhttp.Server.WriteTimeoutof 60 s is deliberately untouched).- Intent bumps (web). Every operator-intent handler bumps the customer's generation AFTER its
successful store write (fire-after-commit, never on an error path), via nil-safe
s.bumpIntent: config create/update/delete, claim resend, offsite re-issue (UI + the re-enroll seam), offsite freeze/unfreeze, retrieval-password regen, block/unblock, per-customer floor, global floor (bumps every config-managed customer), controller log-tail request, and the CONTROLLER log-bundle request (the AGENT ring rides the heartbeat envelope — a separate plane, deliberately not bumped). - Wiring + shutdown. One
intent.New()inmain.go, injected into both the web server and the API handler;intentHub.Close()runs beforeserver.Shutdownso held waits complete instantly instead of eating the 15 s grace window. nil-safe throughout (an unset hub → wait 503, bumps no-op).
v0.57.0 — reinstall-of-existing-customer arc: claim continuity, offsite re-issue, escrow honesty (2026-07-16)
Closes the N100 physical-run findings F2/F3 and the correctness edge behind F4→2.3
(documentation/tests/VALIDATION-n100-baremetal-2026-07-16.md). When an existing customer's box is
clean-slate reinstalled, hub and box previously disagreed about claim, offsite, and escrow state.
This makes the reinstall path first-class — the Peti (R-1) convergence prerequisite.
- F2 — claim continuity.
claim.Engine.ReissueForReenrollrides the existing rotation semantics: for a CLAIMED customer whose box re-enrolls (fresh, passwordless), it bumps the generation ONCE and emails a RESET code (delivery via the existing report ACK) — the customer no longer has to hunt for the manual "request a new code" button. No-op for an unclaimed customer (first-provision path owns the code). Hooked at the host-enroll mint path (handleHostEnroll), which fires exactly once per fresh host record — the single-bump-per-re-enroll guarantee. Emitsclaim_reissued_reenroll. (Fork verdict, source-verified: the hub stores only the claim code + a claimed boolean — never the password hash, which is controller-owned by the arc's design. So fork B, not A.) - F3 — offsite continuity. The re-enroll mint path also calls the same machinery as the manual
"Re-issue offsite credentials" button (
web.Server.ReissueOffsiteForCustomer, wired to the api handler viaSetOffsiteReissuer) — the one-time offsite password only ever reached the OLD controller, so the fresh box gets a fresh one and aConfigVersionbump. Emitsoffsite_reissued. - 2.3 — escrow honesty (correctness; red-proofed). Re-issuing offsite credentials changes the
restic repo password, so any existing key-escrow blob is now STALE (a recovery code minted against
it would decrypt a password that no longer opens the repo).
offsite.ReissueCredentialsnow marks the escrow stale (store.MarkEscrowStale, cleared by the next ceremony viaSaveHostEscrow); the ACK withholds the now-mismatchedrestic_pw_sha256so the controller cannot auto-confirm against a dead key, and the DR-tier checklist shows stale instead of "ceremony done." Emitsescrow_stale. Red-proof: withMarkEscrowStalegutted, the hub keeps advertising ceremony-done after a re-issue →TestReissue_InvalidatesEscrowFAILS; restored → passes. - Out of scope (reported): F4's general installer fix is NOT feasible — the DR storage id lives in
the agent-domain pbs_dr descriptor (provisioned post-WG), not the installer-fetched config, and the
underlying block is the agent's token-auth pre-check 403ing before its own root-run
felhom-pbs-apply grant. Root fix is agent-side (ROADMAP agent-train item); the demo was unblocked live with a one-shot ACL grant. Controller (Part 3) unchanged: its escrow prereqs are already fetched live from the agent, so F4/Part-0 alone restore them (3.1 spec premise contradicted by source). scripts unchanged (v1.16.0).
v0.56.0 — PBS-DR self-heal reconciler (re-stage a consumable secret) (2026-07-15)
Implements SPIKE-pbsdr-selfheal-2026-07-15 (e8f8c44). ⚠️ ARCHITECTURE IMPACT: before this,
there was no automatic recovery for the commonest real event — a customer box re-installed /
restored / rolled back onto its stable host_id. The hub kept the durable pbs_dr descriptor
(enabled) + a durable consumed one-time secret; the WG peer still existed (same pubkey →
changed==false, so the provision cascade could not re-fire) and applyPBSDR's "already
provisioned → no-op" meant even a config re-save minted nothing. The agent sat in
pbs_dr.state="waiting_secret" forever — PBS-DR never converged, so escrow could not run and offsite
never armed. The spike proved (SQ-2b′) the missing piece is a consumable secret, not the
descriptor: re-staging the stored secret converged a stuck box in one ~30 s agent tick, using the
existing ep0 token, zero churn. This reconciler closes that gap.
internal/pbsdrheal/reconciler.go(new): a hub periodic reconciler (5 min,wgsyncshape). For each host whose descriptor is enabled + provisioned and whose latest reportpbs_dr.stateis a stuck state held across a debounce (≥2 distinct reports — so a box brieflywaiting_secretbetween provision and its first consume is not touched):waiting_secret→ re-stage the stored secret (no ep0 call, no generation bump); no stored secret → escalate to Re-issue;consumed_failed→ escalate to Re-issue only (a re-stage would re-feed a burned secret). A converged/disabled/verify_failed/unprovisioned/DR-OFF host is a pure no-op. Each heal emits a distinct audit event (pbsdr_selfheal_restaged/_reissued/_consumed_failed). Reads the hub DB only — never the box.internal/store/pbsdr.go:RestageHostPBSSecret(hostID) (bool, error)— clearsconsumed_atIFF a row exists (no INSERT, no value change, no generation bump);restaged=false→ the caller escalates.PBSDRHealStates()— one query joining each host's descriptor enable/provision flags to its latest report'spbs_dr.state+ id (mirrorsGetHostOOBStates).internal/web/pbsdr.go:ReissuePBSDR(ctx, customerID)— the non-HTTP core of the operator Re-issue button (tenantsync reissue → fresh consume-once secret → descriptor refresh + gen bump), now the reconciler's escalation seam. The operator handler is unchanged.cmd/hub/main.go: the reconciler is started unconditionally (the primary re-stage heal needs no endpoint).PBSDRHEAL_ONLY_HOSTenv scopes a supervised first rollout to one host (empty = whole fleet — the steady state).- Tests:
internal/pbsdrheal/reconciler_test.go(Scenarios A–F + scope + no-re-heal, real store- fake action seam) and
internal/store/pbsdr_test.go(re-stage semantics + the no-generation- bump guard +PBSDRHealStatesparsing). All six §10 red-proofs verified (mutation → FAIL → revert). No agent changes (the agent already self-heals once a secret is consumable).
- fake action seam) and
- Live-leg (drill guest qm300): rolled back to
post_day0_golden136(the re-install reproduction) → stuckwaiting_secret(marker gone,felhom-pbsabsent, hub secret consumed). Deployed scoped viaPBSDRHEAL_ONLY_HOST=demo-vm-felhom-2f4b00. The reconciler observedwaiting_secretacross two reports (16:52 + 17:07 UTC) and at 17:10:00 re-staged the stored secret (eventpbsdr_selfheal_restaged; "no ep0 token minted, no generation bump"); the agent re-consumed at 17:10:26 and converged (state=applied) at 17:10:28 — hands-free, no operator click. The demo hostdemo-felhom-01was never touched (out of the scoped work set). Fleet-wide widening (removePBSDRHEAL_ONLY_HOST) is a deliberate operator follow-up — not done in this supervised session.
v0.55.0 — accept the offbox_enlarge_blocked event (Task 3a-fix delivery chain) (2026-07-15)
The controller (v0.134.1) sends an offbox_enlarge_blocked warning when an app's enlarged offsite
push is refused by the quota gate (config+DB still saved). This closes the ingestion link of its
delivery chain.
internal/api/handler.go:offbox_enlarge_blockedadded toallowedEventTypes— the event was 400-rejected before (the customer email was dropped at ingestion). Acceptance case + 400-red-proof added tointernal/api/event_test.go.- Deliberate NON-change: NO
customerMessagesentry.FormatCustomerEmail(templates.go:129) gives the static per-type message PRIORITY over the raw message, so a static entry would DISCARD the controller's dynamic two-number Hungarian text (estimate + quota). The documented fallback (raw message survives) is the correct path; locked by a newinternal/notify/templates_offbox_test.goassertion.internal/notify/dispatcher.goand the customerEnabledEventswhitelist are unchanged here — the controller side (v0.134.1) ownsDefaultEnabledEvents+ the prefs migration + the settings checkbox.
v0.54.0 — operator login password changeable from the UI (2026-07-13)
The hub login password was previously settable ONLY by editing the auth.password_hash field in
the hub-config ConfigMap and redeploying — there was no in-app way to change it. Added a
"Login password" card on the Configuration page.
- DB-override precedence (same pattern as the controller-version floor). New
hub_settingskeyoperator_password_hash(store:Get/SetOperatorPasswordHash). The web server no longer reads a static field for auth:Server.passwordHashis renamedconfigPasswordHash(the hub.yaml SEED) and every auth check — the CSRF gate,RequireAuthsession/basic-auth paths, andhandleLogin— now goes througheffectivePasswordHash()= DB override wins, else config seed. The ConfigMap value stays the break-glass fallback: blank the DB row (or edit the manifest + redeploy) to reset a lost password. POST /configuration/password(handleChangePassword): requires the current password (verified against the effective hash), a new password of 8–72 bytes, and a matching confirmation; rejects a no-op change. On success it bcrypts the new password (cost 10, matching the seed) and persists the override. Existing sessions are intentionally kept valid — only the next sign-in and Basic-Auth use the new hash. CSRF-enforced (existingServeHTTPgate); no secret is ever logged.- UI: change-password card on
configuration.htmlwith current/new/confirm fields, inline client-side mismatch pre-check, and six flash outcomes (pw_changed,pw_current_wrong,pw_too_short,pw_too_long,pw_mismatch,pw_unchanged). - Tests + red-proofs (
change_password_test.go): override-wins precedence, happy-path end-to-end throughhandleLogin(new works, old dead), wrong-current rejection (security anchor), mismatch/too-short/no-op rejections, and template render. Red-proofs verified — dropping the current-password check writes the override anyway (WrongCurrentRejected fails); breaking the override precedence kills both the precedence and happy-path login assertions.
(unreleased) hostInstallVersion 1.16.0 (2026-07-13)
Display-const bump only, keeping scripts/hostinstall_gates.py green with the installer's
v1.16.0 (FELHOM_ESCROW via the canonical sudoers fetch — see scripts/CHANGELOG.md). No behavior
change; rides the next hub image train (no deploy for this).
v0.53.0 — closing bundle: F-14 gated auto-Reissue + dead-host roll-up honesty + bearer out of git (2026-07-13)
The last engineering items on the pre-tester board. Two operator rulings in force (CONTEXT.md): F-14 auto-re-issue only on a recorded escrow-acked deletion; customer status never better than its worst expected host.
- F-14 deletion provenance + gated auto-Reissue (take-two MEDIUM: host delete + re-enroll
with a surviving ep0 tenancy = DR re-attach dead-end,
token_existson both auto-provision and config save). Newhost_deletionstable — host_id, customer_id, deleted_at,escrow_acked(= ack given over a PRESENT escrow row) — written INSIDE the DeleteHost transaction; NO backfill (pre-record deletions keep the manual path by design). The provision atom, ontoken_exists, reads the customer's MOST RECENT deletion record: escrow_acked → invoke the EXISTING tenantsync Reissue op, store thepbsdr_auto_reissueaudit event ("Previous key destroyed (acknowledged deletion) — credentials re-issued automatically."), proceed; no record / un-acked → the pre-existing refusal byte-unchanged (never-silently-re-key law). Red-proofs: provenance-write drop → scenario-A fails; gate bypass → scenario-B's zero-reissue assertions fail (silent re-key visible as a 303). - Dead-host roll-up honesty (drill-1 observation, live on the Peti cluster: proxmox1 down
23h behind a GREEN customer row — controller reports ride the internet, independent of the
agent). Customer status on the dashboard, /configs list and customer detail (header + strip)
is now
worst(controllerDerived, hostStatusOf(each expected host))via THE single staleness definition (Server.hostStatus; no second threshold anywhere): any host down/stale caps the customer at WARN with a cause chip naming the host ("host down: "); pending hosts worsen only after the customer has ever reported (onboarding exclusion). The three inlined controller-status chains collapsed intocontrollerStatus()(rollup.go). Display + derivation only — HostStalenessChecker alerting untouched. Red-proof: fold removal → the exact Peti fixture renders green → TestRollup_DeadHostMasking fails. - Operator bearer out of git (the two publish runbooks' ROTATION item):
manifests/hub.yamlno longer commitsreport_api_key— the Deployment injectsREPORT_API_KEYfrom out-of-bandSecret/report-api(deliberately NOToptional:— a missing Secret fails Ready instead of booting an unauthenticatable hub); main.go gains the env override (RESEND_API_KEY twin). New gatescripts/manifest_bearer_gate.pyblocks bearer-shaped (64-hex) literals in manifests/ (red-proven: reintroduction → exit 1). The controller repo's example-config copy of the literal is scrubbed. The exposed git-history value dies with the SUPERVISED rotation — procedure + full consumer list in documentation/runbooks/secrets.md §"Operator/global bearer key" (per-customer/per-host keys unaffected).
Hub half of the polish batch (take-two findings F-15/F-16). Companion: controller v0.123.0.
- F-15 instant reset codes:
POST /api/v1/claim/reset-requestnow returns the ACTIVE code state in the response —{claim: {code_hash, generation, issued_at}}, the exact shape and bcrypt-only guarantee of the report ACK — so the box applies the rotated hash in the same request cycle and the emailed code works immediately (previously the box learned it only on its next report ACK, ~15 min — Viktor's take-two live failure). Served on every authorized outcome (a cap-reached refusal returns the unrotated row = controller-side no-op by generation). Never a plaintext code. Live-proven: apply 1 s after the request; code accepted on first try. The operator "Kód újraküldése" (no box round-trip) still has the ACK lag — its flash + data-confirm copy now say so ("A kód a doboz következő jelentésekor (~15 percen belül) aktiválódik."); the previously unmappedclaim-resent/claim-resend-failedflashes render. - F-16 inline confirms: every native
confirm()in the hub UI (offsite re-issue, freeze/ unfreeze, PBS re-issue, telemetry reset, dismiss-all-issues, regen-password, claim-resend, block/delete, geo-disable) replaced by the LIGHT inline two-step — the sharedinline_confirm.htmlpartial (felhomConfirm+data-confirmdelegation,requestSubmitso formaction survives). Native confirms are OS-modals that froze CDP browser automation (F-16; the drill F-11 siblings). NOT the danger-zone typed-confirm — that heavyweight cascade flow is untouched. New gatescripts/hub_confirm_gate.py(zero native confirm/prompt; red-proven). Live-proven: the offsite re-issue completed under automation without freezing.
v0.51.0 — DR-tier-by-default: per-customer flag + hands-free cascade + offsite coupling + capability chips (2026-07-12)
Hub half of the DR-tier-by-default batch (DRILL-day0-vm-2026-07-12; operator decisions 1–5: capability BAKED on every install, activation is THIS flag, DR defaults ON for new customers, identity-only escrow PARKED by policy, WG is base infrastructure). Companion: installer v1.15.0
- agent v0.86.0 (capability
inactivestate).
- Per-customer
dr_tierflag (customer_configs column + form checkbox in the renamed "DR tier (PBS, ep0)" section, replacing the oldpbsdr_enabledform field). NEW customers default ON; legacy rows were initialized FROM REALITY by a one-time migration backfill (host carries an enabled pbs_dr descriptor → ON, else OFF — never auto-cascade a legacy box; backfill runs only on the ALTER that adds the column, so later operator opt-outs survive). - Cascade semantics (scenario D): an UNMET precondition (no host / no WG peer / no tenantsync) is no longer a save-blocking error — the flag stores the intent and the edit form shows per-stage status (host enrolled → WG peer → descriptor provisioned → escrow present), reusing the fail-closed guard wording. REAL provisioning failures stay fail-closed (tenantsync error, token-exists → Re-issue).
- Hands-free auto-provision (scenario A): a host's FIRST WG peer registration fires
PBSDRAutoProvision(apiSetWGRegisteredHook, wired when tenantsync is enabled) — a DR-ON customer's descriptor provisions with ZERO operator steps and applies on the agent's next desired-state tick. Detached goroutine; never delays/fails the registration response. - Offsite requires the DR tier (scenario C, drill F-6 CLOSED BY POLICY):
applyOffsiterefuses without the flag — exact message "Offsite backup requires the DR tier — enable it first (the escrow ceremony depends on the PBS key)". No more provisioning into the EscrowState-pending-forever dead end. - Capability chips on the host page (NEW render surface): the agent's privileged-capability
self-check is now visible — ok (blue), degraded (warn / error when critical), and the agent
v0.86.0
inactivestate as a NEUTRAL chip (disabled ≠ degraded). The pre-v1.15.0 pbsdr "binary not found" signature surfaces the migration one-liner (never silently pretend)..badge-okfinally defined in style.css (was referenced, fell back to bare.badge). - Setup-tab installer copy:
hostInstallVersion1.12.0 → 1.15.0 (drill F-1), now gated against the installer's SCRIPT_VERSION byscripts/hostinstall_gates.py. - Tests + red-proofs (all four mutations proven red): coupling gate (guard removed → refused case fails), flag default (default flipped → form test fails), backfill (enabled:false ignored → disabled case fails), auto-provision (hook unhooked → scenario A test fails); plus cascade-wait, chips render (inactive-neutral / degraded-stays / migration hint), and the one-time-backfill-survives-reopen case.
v0.50.0 — customer-claim password arc: code engine + email + ACK/config delivery (2026-07-12)
Hub half of the customer-claim password gate (closes DRILL-day0-vm F-4/F-5; needs controller
v0.122.0). The customer OWNS the dashboard password — the hub generates a one-time claim code,
emails it (Hungarian) to the REGISTERED address, and stores only bcrypt(code). No operator-set
path; the plaintext code exists solely inside the email send (the retrieval-passphrase custody
rule).
internal/claim— the code engine.EnsureIssued(idempotent — issue+email at the FIRST real config retrieve = Day-0, and at a live box's first report; repeated pulls/reports never rotate or re-send),Resend(operator button; rotates generation — unclaimed gets the claim template, claimed gets the reset template),RequestReset(controller-forwarded "Elfelejtett jelszó", rate-limited 3/day/customer),MarkClaimed(set-only; one confirmation email on the unclaimed→claimed transition).store.customer_claims— per-customer{code_hash, generation, issued_at, emailed_at, claimed_at, reset_day, reset_count}.RotateClaimCodebumps the generation (single active code) and PRESERVESclaimed_at(a reset never un-claims);MarkClaimedis set-only.- Delivery:
GET /api/v1/config/{id}bakesweb.claim_code_{hash,generation,issued_at}into the generated controller.yaml (gate-from-first-boot) and issues the first code; the report ACK serves the activeclaim {code_hash, generation, issued_at}(allowlisted) and ingests the controller'sclaimedflag (set-only).POST /api/v1/claim/reset-request(self-scoped by the box's report key). New emails via the notify dispatcher;claim_lockoutevent allowlisted. - UI: the customer page Setup tab shows a claim status chip (Nyitott — kód kiküldve / Claimed)
- a "Kód újraküldése" button (
POST /configs/{id}/claim-resend) — no plaintext code ever rendered (there is none to render).
- a "Kód újraküldése" button (
- 15 tests (engine, ACK/config, UI); the arc's red-proofs live in the controller repo (gate) + here (generation bump, reset non-DoS).
v0.49.0 — Edit tab merge (edit-a), scoped auto-refresh, style.css cache-bust (2026-07-12)
The task spec targeted "v0.48.0", but v0.48.0 (app_start_failed, below) had already shipped + deployed by the time this train ran — a published tag is never re-pointed, so this is v0.49.0. Baseline
3e949bc; commitse740147→2e03de1→1d94b1a→ docs/manifest.
- Edit tab merge (edit-a) (
templates/customer_unified.html,templates/config_form.html, newtemplates/config_form_body.html,web/configs.go,web/pbsdr.go): the standalone customer edit page merged into the customer page's Settings tab, renamed Edit. The form body is a shared{{define "config_form_body"}}sub-template (thehost_detail_bodypattern) built by the oneconfigFormDataview-model builder; the standalone chrome keeps rendering it for the create flow (/configs/new) and the validation-error re-render. The Edit tab renders: config form, Controller Update card, Geo card, and a Danger zone card holding the Block/Unblock/Delete forms relocated verbatim from the Customer Info header (endpoints +confirm()unchanged) — all SIBLINGS after</form>(nested forms are invalid HTML and would break the offsite/PBSformactionsub-buttons). The header keeps only the config-less Create Config action.GET /configs/{id}/edit→ 302/customers/{id}#tab=edit; tabs JS gains thesettings→editlegacy-hash alias. - Server-side required fields on update (
web/configs.gohandleConfigUpdate): the twin of the form'srequiredattributes (Display Name + Domain), checked BEFORE provisioning; the error path re-renders the standalone page with the SUBMITTED overrides so typed values are never lost (red-proofed: nil overrides → values reset → test fails). - Redirect anchors: update/block/unblock/offsite-reissue/offsite-freeze/pbsdr-reissue →
?flash=…#tab=edit; regen-password →#tab=setup(its card lives there); delete unchanged (/configs?flash=deleted). - Scoped auto-refresh (
templates/customer_unified.html): the 60s reload fires only while a live tab (data-live-tabs="overview,applications,events,host"on the nav) is active AND no form is dirty (delegated document-level input/change listener, never reset — a reload clears it). Skipped ticks reschedule; a muted(paused)hint shows next to the toggle on non-live tabs / dirty forms. Toggle,hub_auto_refreshlocalStorage key, cadence, default-on: unchanged. - style.css cache-bust (all
templates/*.html): every stylesheet link is now/style.css?v={{hubVersion}}— closes the v0.47.0 gotcha (max-age=3600served stale styling for up to an hour after each deploy). Red-proofed (bare link fails the render test). - Repo staging rule (
CLAUDE.md): nevergit add -Ain this repo (the v0.47.0146d165sweep incident) — explicit paths, pull-rebase, one writing session per clone. - Tests: +11 (Group A panel surface / sibling-form / header-count, Group B redirect + create +
typed-values-preservation table, Group C refresh structural pins, Group D cache-bust sweep).
Amended pins:
customer_tabs_test.go(settings→edit),pbsdr_test.go(postUpdate supplies the now-required fields; FormRendersState asserts the embedded Edit-tab render).
v0.48.0 — accept the app_start_failed event (controller fix-3, CAMPAIGN-3) (2026-07-12)
app_start_failedadded toallowedEventTypes(internal/api/handler.go) +customerMessages(internal/notify/templates.go). Without the allowlist entry the controller's fix-3 event (a DEPLOYED app found not running — controller v0.120.0) would 400 at ingest and never reach the operator. No other hub change; pairs with controller v0.120.0 which closes the CAMPAIGN-3 finding set.
v0.47.0 — UI reorganization: customer tabs, Host tab, stale-host removal, offsite multi-endpoint UI, button contrast (2026-07-11)
Five hub-side deliverables; no agent/controller/protocol changes. Baseline 8e1a3f0
(v0.46.0); commits 9f29bf3 → ae950e5 → 146d165(swept WIP) → 068427a → 0daddcd.
- CSS button contrast (
templates/style.css):.data-table td a→:not(.btn)(base + hover) —<a class="btn">inside data-table cells (host-detail Diagnostics View/Download, customer log-tail buttons) rendered blue-bright on blue-bright, i.e. invisible. Plain table links keep the bright-link style;.btnitself untouched, no!important. - Customer page tabs (
templates/customer_unified.html,style.css): the ~18 stacked sections split into 8 client-side hash tabs (#tab=overview / applications / setup / settings / backup / events / notifications / host) + a sticky summary strip (name, status, controller version, last report, containers chip). Graceful degradation is load-bearing: panels hide only under a JS-addedbody.js-tabsclass — no JS = every section visible, all existing render tests pass unmodified. Events tab carries a red error-count badge (reuses the already-fetchedCountEventsBySeveritydata — no new query). The auto-refresh reload preserves the hash → the active tab survives. No handler/data-model change for the tabs. - Host tab + shared sub-template (
templates/host_detail_body.html,web/hosts.go,web/configs.go,store.ListHostsByCustomer): the host-detail body extracted into a{{define "host_detail_body"}}rendered by BOTH/hosts/{id}(chrome + call) and the new per-customer Host tab (a LIST by design — 1 host today, N for a later HA cluster; empty state otherwise).handleHostDetail's data assembly extracted intohostDetailData. - Stale host removal (
store.CountHostArtifacts/DeleteHost,web/hosts.gohandlers, routes above the/hosts/catch-all):GET /hosts/{id}/delete-impact(counts/booleans ONLY) +POST /hosts/{id}/deletebehind a type-to-confirm dialog (global-floor pattern). Gates: ONLINE host → 409 always (no override — a live agent would 401 forever; enroll is passphrase-gated mint-once); confirm mismatch → 400; escrow present without the explicit checkbox → 409 with the tx never started (ErrHostEscrowPresent, fail-safe-to-refuse). One transaction cascades guests, host_reports, signed_jobs, host_recovery, host_pbs_secrets, host-scoped log bundles (scope_id == host_idONLY — customer-scoped bundles survive), the bound wg peer (inside the tx — no stranded peer on crash), escrow (only when acked), then the host row. The wgsync 5-min declarative push converges the endpoint afterwards — no reconciler change. Danger-zone card renders only when deletable, so the hosts-list zero-<button>pin and the detail-page 2-button pin stay green unmodified. - Offsite multi-endpoint UI (
store/wg.goListWGEndpoints/DeleteWGEndpoint,web/offsite.go,templates/offsite.html):/offsitelists ALLwg_endpointsrows as cards + add/edit/delete forms (posture change from S2 read-only — operator decision). Validation → 400 stores nothing; subnet edit / endpoint delete refused 409 while peers sit in the (current) subnet; pubkey change gets a type-to-confirm noting pull-based convergence. Peer table gains an Endpoint column (first id-ordered subnet match, em dash when none). Allocation, reconciler push and desired-state merge stay lowest-endpoint-id (GetWGEndpointuntouched; the page states the deferral) — per-endpoint allocation (wg_peers.endpoint_idmigration) is a future arc. - Tests: +21 new/amended across web+store: tab render (no-JS completeness, badge, banner-
above-tabs), Host tab shared-body/empty/isolation, delete cascade + refusal non-effects +
impact shape + danger-card gating, offsite cards/column/validation/guards. Five companion
red-proofs ran and FAILED as required (online gate, escrow ack, bundle scope, endpoint
delete guard, subnet-change guard). Note: commit
146d165(parallel session) swept the Part-4 WIP mid-red-proof —068427arestored the escrow-ack line.
v0.46.0 — observability pass: per-box log pulls, bundle custody, TTL + secret gate (2026-07-11)
Hub third of the cross-repo observability task (agent v0.83.0 + controller v0.116.0): remote, pull-only access to both box components' always-DEBUG capture rings — honest to the sovereignty posture (the hub never connects in; the box pushes on its own cycle and its own log records the pull, customer-visible).
- Store (
internal/store/logbundle.go+ schema):log_bundle_requests(one pending intent per scope+component; scope = customer_id for the controller/report channel, host_id for the agent/ heartbeat channel) +log_bundles(gzip payload, newest-3 retention, 72 h TTL purged on the existing 60 s sweep).SaveLogBundleruns the token-pattern secret gate BEFORE storing — a hit stores ablocked: possible secretflag row with NO payload (fail-closed; WARN logged hub-side);[REDACTED]shapes and public checksums deliberately pass (red-proof: gate disabled → the plantedre_…token stores → FAIL). - Channels (additive, both directions): the report ACK gains
controller_log_requestedand ingestscontroller_log_tail; the heartbeat envelope gainslog_tail_requestedand ingestslog_tail. Consume-once on arrival (red-proof: clear-on-arrival removed → the ACK re-advertises forever → round-trip tests FAIL). A pre-0.83 agent simply never fulfills — the request stays visiblypending(S6 tested), harmless. - UI (host detail, English like the rest of the hub operator surface): a Diagnostics section
with Request controller logs / Request agent logs buttons (CSRF form posts), state rows
(
pendingwith the honest per-channel latency hint — controller ≤ ~15 min report interval, agent ≈ heartbeat cadence — /availablewith View+Download /blocked), and the 72 h custody note. The hosts read-only invariant is amended: these two request forms are the ONLY actions (pinned by test). - Download endpoint
/hosts/{id}/log-bundles/{bid}[?download=1]— session-authed like the rest of the operator UI, scoped to the host's own channel scopes.
v0.45.0 — floor-UI separation + effective-floor source + per-box MinAgent conditional floor (2026-07-11)
Two parts of the NAS/coupling backlog, both addressing the publish-train 0.81/0.113 floor footguns.
- Floor-UI separation + confirm (Part C): the global controller-version floor is its own card with
a type-to-confirm dialog that first states the live blast radius —
GET /configuration/global-floor/impact?v=X.Y.Z(countBoxesBelowFloor, honoring per-customer overrides) — so "the floor acts immediately" is impossible to miss. An effective-floor + source line (store.ResolveGlobalFloor→GlobalFloorResolution) shows the resolved value and WHICH source won (DBhub_settingsvs envDEFAULT_MIN_CONTROLLER_VERSION, both raw values when they differ) — the 9-minute-skew incident's root cause, now permanently visible. The Day-0 artifact manifest save is unchanged and provably does NOT touch the floor (regression-tested). - Per-box MinAgent conditional floor (Part D): the artifact manifest gains
MinAgent(the golden's controllerMinAgent:header; blank = uncoupled, no gating). At report-ACK timestore.ResolveManagedFloor(customerID)compares the box'shosts.agent_versionagainst it: agent ≥ MinAgent → the controller floor is served; below or unknown → the floor is HELD (ACK omits the directive) and the box is flagged on the Hosts dashboard (floor held: agent <v> < MinAgent <w>). Mechanises the "agent BEFORE controller floor" rule per box — the manual fleet check is retired (publish-train-rules.md rule 3 updated). - THE one comparator:
web.compareVersions's body moves to a leafinternal/semverpackage (Compare/Valid); web delegates, store's MinAgent gate reuses it (no import cycle, no second comparator; gitea's documented local copy is out of scope). - Tests + red-proofs: floor source precedence (DB-wins), manifest-save-doesn't-touch-floor, impact-count with override exclusion, source-line render; managed-floor hold/serve/uncoupled/ unknown-agent + a fleet discriminator + the report-ACK wire test (held box omits the floor, served once the agent qualifies). Every red-proof run → predicted failure → reverted.
v0.44.0 — PBS DR tier SLICE 1: ep0 tenantsync surface + hub provisioning (2026-07-10)
Builds on SPIKE-pbs-tier-provisioning (00afadc). The operator ticks "PBS DR tier (ep0)" on a
customer config → the hub verifies the host's WG peer (the agent self-registers it; absence is
fail-closed) → provisions the per-customer ep0 PBS namespace + privilege-separated token over the
NEW felhom-tenantsync forced-command surface (peersync untouched) → stores the token secret
CONSUME-ONCE, host-scoped → serves the non-secret descriptor via the host desired-state (generation
bump). The agent apply-bridge is SLICE 2 — nothing is live-provisioned yet.
scripts/felhom-tenantsync.shv1.0.0 (installed on ep0 per runbook §10): JSON-on-stdin/stdout; opsprovision(existing token = hard errortoken_exists— re-issue is explicit),reissue(delete-token purges ACLs → recreate → re-grant),fingerprint. Dual-grant per spike §3; own-namespace self-check with one regen retry (the spike's transient-403 note) then rollback. Secret hygiene: the token secret rides stdout ONLY (all tool stdout → stderr; never a file/argv). NO deprovision op — namespace/data deletion stays a deliberate, separate decision.internal/tenantsync: the wgsync twin — pinned host key (exact-match, constrained HostKeyAlgorithms), per-op JSON exec, typedErrTokenExists. Divergence from wgsync: error messages NEVER embed stdout (the secret channel) — red-proof-style contract test (TestErrors_NeverEmbedStdout).- Store:
host_pbs_secrets(host-scoped consume-once, theone_time_secretstwin) +SaveHostPBSSecret/ConsumeHostPBSSecret(same-tx mark; re-save resets). - API:
POST /api/v1/hosts/{id}/pbs/consume-token— per-host key, self-scoped (global key = operator recovery); 200 exactly once → 404; a foreign key's 403 does NOT burn the secret. (Task spec wrote/host/{id}/…; implemented under/hosts/for namespace consistency with every other agent-facing route.) - Web: config-form section "PBS DR tier (ep0)" (enable + storage-id, default
felhom-pbs— the descriptor carries the id so the slice-2 bridge is name-agnostic and the demo'sfelhom-offsiteadoption dissolves the naming collision) + provisioned line + Re-issue PBS credentials (the offsite F4 precedent).applyPBSDRmerges the non-secretpbs_drdescriptor into the HOSTdesired_json(admin-set path,SetHostDesiredbump) — ConfigJSON never carries it. Fail-closed on: no tenantsync key, no enrolled host, no WG peer, no endpoint record, tenantsync error,token_exists(message points at Re-issue). Already-provisioned re-save = success-no-op (no re-key, no second secret, no spurious bump). Disable = descriptorenabled:false, tenancy kept. - Deploy:
manifests/hub.yamlgainsTENANTSYNC_SSH_KEY_FILE+ optionalSecret/tenantsyncmount (same endpoint addr + pinned host key as peersync, its own key). - Red-proofs run and recorded (REPORT.md): consume-once mark drop → the secret re-serves (store + API layers); fail-closed guard swallow → 303 half-save; idempotency short-circuit drop → token rotation + fresh secret + spurious bump.
v0.43.1 — Git Sync form hint: credentials are optional (2026-07-10)
Pairs with controller v0.112.0 (anonymous registry self-update). The config editor's Git Sync section
looked load-bearing; in truth the credentials matter only for a private app catalog — version discovery
and self-update work without them since controller v0.112.0. One template hint added
(config_form.html): "Opcionális — csak privát alkalmazás-katalógushoz. A verziófrissítés enélkül is
működik." No behavior change.
v0.43.0 — Remote app-log diagnostics: copyable issues + error context + on-demand log tails (2026-07-10)
Pairs with controller v0.111.0. Motivated live: Peti's CWA NFS issue was visible in Known Issues but tooltip-only unreadable and context-free, and there was no way to see the app's actual logs without box access. The hub still never connects into a guest — everything rides the existing report + ACK.
- Part A — readable, copyable issues (
templates/app_detail.html): Known Issues rows are click-to-expand — full message in a wrapping monospace<pre>+ Copy button (clipboard API with execCommand fallback), fingerprint/severity/first-last-seen in the body. Tooltip-only truncation killed. - Part C — context stored + rendered (
store/telemetry.go):app_log_issuesgainscontext(JSON array) +context_customer(provenance);upsertAppIssuestores context on INSERT and adopts a later one ONLY while the stored context is empty (first capture wins — stable repro, no churn). Rendered in the expanded row as "Context around first occurrence — from ", copyable. Nil-safe with pre-v0.111 reports. - Part D — on-demand ordered log tail (pull-based): per-app "Request log tail" button on the
customer page →
log_tail_requestsrow (one active per app; re-click refreshes) + a customer-visiblelog_tail_requestedevent (transparency by default). The report ACK advertiseslog_tail_requests: [app…](same additive omit-when-empty pattern as escrow); the controller's next report shipslog_tails→ stored inapp_log_tails(transient, last 2 per app kept) and the request is cleared (consume-once). Ordered tail view with line numbers (log_tail.html) + Download .log; tail reads are customer-scoped. - Part F fix — the 24h/7d/30d selector now filters Known Issues:
GetAppIssuesgained the samesincecutoff the Memory Trend uses (it had NO time filter — a 24h view showed 25-day-old rows). - Part G — deletion → dismissal: diagnosis = the delete handler was NOT broken; deletion is futile
because the controller re-scans its rolling 15-minute window every report and re-upserts a still-
occurring fingerprint with fresh
last_seenminutes later. Replaced withdismissed_at: Dismiss Selected/All (buttons renamed), dismissed rows out of the default view ("Show dismissed" toggle), andupsertAppIssueun-dismisses ONLY onexcluded.last_seen > dismissed_at— a re-sent old window stays hidden, a genuinely NEW occurrence resurfaces (recurrence never silently swallowed). - Part H — per-customer scoping:
?customer=<id>on the app detail page filters Known Issues to rows whoseaffected_customerscontains the id (header shows "filtered: "); the customer page's App Telemetry rows link there (the drill-down). The fleet view stays; the expanded row lists the affected customers explicitly (linked) and the count column is labeled "Occurrences (all customers)". - Tests + red-proofs (all four failed exactly as designed, restored green): dismissal guard dropped → old-window re-report resurrected the row → FAIL; range predicate neutered → 10d-old issue visible at 24h → FAIL; first-capture-wins dropped → empty-context upsert clobbered stored context → FAIL; consume-once DELETE removed → request survived fulfillment (store test + API ACK round-trip both) → FAIL. Plus: late-context adoption, warn-no-context, occurrence counting, tail request/fulfill/prune- to-2/cross-customer-404, ACK omit-when-empty baseline, render tests (expanded row content, customer page sections, ordered tail view + download headers).
v0.42.0 — Remote "Debug mód" toggle on the customer config editor (2026-07-10)
Lets an operator flip the controller's debug mode (verbose log + the /debug menu, which the controller
gates on Logging.Level=="debug" / isDebug()) remotely, without SSH — the support workflow (today:
Peti's box). The config-version bump on save makes the controller re-pull + self-restart on its next
report ACK, so the switch takes effect within a cycle, hands-free.
- Form field, not raw-JSON injection — on purpose.
handleConfigUpdateREBUILDSConfigJSONfrom the form on every save (buildConfigJSON), so any foreign key injected straight into the stored JSON is dropped on the next save. The toggle is therefore a real form field, which by definition survives every save. (The offsite descriptor survives via its own separate provision-merge, untouched by this.) buildConfigJSON(internal/web/configs.go):debug_modechecked → emits"logging":{"level":"debug"}; unchecked → theloggingkey is OMITTED entirely (the generatedcontroller.yamldefault stands — no needless"info").- Config form (
templates/config_form.html): new collapsible "Hibakeresési mód (fejlesztői)" section with thedebug_modecheckbox; render state parsed back fromConfigJSON(logging.level=="debug"→ checked). No change to the offsite/CF/git leg; no generic raw-JSON editor (deliberately — validated surfaces only). - Tests (non-hollow,
configs_debug_test.go): form→JSON both ways (checked emits / unchecked omits); full-path survival test throughhandleConfigUpdateproving the debug key lands, the offsite descriptor is byte-for-byte unchanged across save+re-provision, and the red-proof that a hand-injected foreign key is gone after one save (why the switch must be a form field); render state both ways. Red-proof exercised (feature disabled → survival + form tests fail).
v0.41.0 — SLICE 4: OffsiteChecker (fill + staleness) + operator freeze lever (2026-07-09)
The last build item of the offsite arc (pairs with controller v0.109.0's soft-quota gate + report object).
internal/monitor.OffsiteChecker— a SIBLING of StorageFillChecker (same born/persistent, escalation-only, recovery-re-arm shape; NOT bolted onto the disk checkers), reading the controller report's newoffsiteobject. Two signals: fill (repo_size_bytesvsquota_gbat warn 90 / crit 95 — quota 0 = dedicated, never alerts) and staleness (offsite_stale, warning): enabled+escrowed but no run in >48h (or never) — the silently-STUCK detector; a recently-FAILING offsite is not stale (backup_failedowns that), and pending/disabled targets never alert (normal onboarding — companion red-proof: dropped the escrowed-only filter → the pending customer alerted → test FAILED). Nil-safe on reports without the object (pre-v0.109 controllers). Tie-guard: duplicate same-second latest reports are processed once per sweep. Same 60s sweep as the other checkers.- Freeze lever (operator, MANUAL only):
Provisioner.SetOffsiteFrozen— flips ONLYreadonlyon the exactly-1 labelled sub-account viaUpdateSubaccountAccess(SSH stays on; ambiguity refuses — tested), wired to confirm-gated Freeze/Unfreeze offsite buttons next to Re-issue (shared model only; dedicated is Hetzner-enforced). NEVER automatic — freezing also blocks prune, the customer's only way DOWN from over-quota. RoutePOST /configs/{id}/offsite-freeze(unfreeze=1reverses); action logged, value-free.
v0.40.0 — SLICE 3: store the escrow password-hash + serve escrow status in the report ACK (2026-07-09)
The hub-verified escrow auto-confirm chain, hub third (pairs with agent v0.79.0 + controller v0.108.0). The controller must verify the RIGHT fact — not "a blob exists" but "the blob covers the CURRENT repo password" — so the hub records WHICH password each escrow covers, as a non-reversible sha256 (a 256-bit random secret's hash is safe to store/serve; the password itself never reaches the hub).
internal/store: additive migrationALTER TABLE host_escrow ADD COLUMN restic_pw_sha256 TEXT(NULL on legacy rows — e.g. the demo's — which therefore never auto-confirm; the deprecated manual confirm covers them).HostEscrow.ResticPwSHA256+SaveHostEscrowgains the param (last-write-wins); NULL-safe reads via COALESCE. NewGetEscrowStatusForCustomer(hosts⋈host_escrow; latest-updated wins).internal/api:escrowUploadRequest.restic_pw_sha256,omitempty(the agent emit struct's mirror —TestEscrowUploadContractupdated in lockstep with the agent's half); stored on upload. The report ACK gainsescrow: {identity_blob_present, restic_pw_sha256, created_at}— omitted entirely when the customer has no escrow row (a fresh customer stays pending silently).- Tests: hash stored + legacy-upload reads back NULL-safe as ""; ACK carries the object / omits it without a row; contract mirror.
v0.39.0 — offsite hardening: F4 credential re-issue + F2 scan retry + F5 save UX (2026-07-09)
Part of the offsite-provisioning hardening bundle (pairs with controller v0.107.0 + agent v0.78.0); the
sharp edges from the live e2e (documentation/audits/VALIDATION-offsite-provisioning-e2e-2026-07-09.md).
- F4 (pilot-gating) — "Re-issue offsite credentials":
Provisioner.ReissueCredentials— the EXPLICIT operator recovery for a consumed-password dead-end (fresh-guest DR; consumed-but-failed install). Resets the customer's sub-account password (ResetSubaccountPassword) or dedicated-box password (newResetBoxPasswordinhetznerapi, client+interface+fake) → stores a FRESH one-time secret → the handler re-saves the config unchanged soConfigVersionbumps and the stuck guest's next refresh re-runs the bridge. Hard-scoped: targets ONLY the resource labelledfelhom-customer=<id>; refuses unless the label lookup finds exactly 1 (ambiguity = refuse, no reset, no secret) + companion red-proof (dropped the exactly-1 guard → ambiguous lookup proceeded → test FAILED). NOT implicit rotation —ProvisionOffsitenever calls it. UI: a confirm-gated button on the config form (shown only when provisioned), routePOST /configs/{id}/offsite-reissue(CSRF rides the parent form). The password value is never logged. - F2 — host-key scan retry-with-backoff: a fresh sub-account's DNS lags creation, so the FIRST save
502'd (
no such host, live).scanWithRetryretries on failure (default ladder 2/4/8/16/30s ≈ 60s total, inside applyOffsite's 3-min detached ctx; ctx-abortable; fail-closed past the budget) + companion red-proof (disabled the retry loop → DNS-lag save failed → test FAILED).Provisioner.ScanBackoffinjectable for tests. - F5 — save UX: the config form disables its submit buttons and shows an in-flight notice on submit (the ~25–60s spinner-less save was the re-click bait that caused F1 live).
v0.38.1 — offsite provisioning: detach from the client's request context (live finding F1) (2026-07-09)
Found in the first supervised live run: the offsite save takes ~25s (create + wait + host-key scan) with no
UI feedback, the operator re-clicked, the browser abandoned the first request, and r.Context() was canceled
between CreateSubaccount and SaveOneTimeSecret — the sub-account was created on Hetzner but its
one-time password was lost forever (the controller's consume 404s permanently; stranded resource).
internal/web.applyOffsite: provisioning now runs oncontext.WithoutCancel(r.Context())with a 3-minute absolute timeout — once the create starts, the create→wait→store atom runs to completion even if the client disconnects. Fail-closed behavior unchanged (an actual provisioning error still 502s and saves nothing).- Test
TestApplyOffsite_ClientDisconnectMidProvision(a ctx-honoring fake cancels the request context mid-create): the one-time password must reach the store and the descriptor must merge despite the disconnect. Companion red-proof: reverted to the raw request ctx → the exact live error (subaccount create action: context canceled) → test FAILED. Restored. - Known residuals (recorded, not fixed here): the form has no in-flight spinner/disable (the re-click bait), and a concurrent save can still hit Hetzner's box-level HTTP 423 action lock (surfaces as the fail-closed 502).
v0.38.0 — offsite provisioning SLICE 2 (hub side): capture the box host-key fingerprint (2026-07-09)
Pairs with controller v0.106.0. So the controller can VERIFY the box identity instead of blind-TOFU, the hub captures the box's SSH host-key fingerprint at provision and serves it in the descriptor.
internal/offsite:Descriptor.HostFingerprint(SHA256:…, non-secret).ProvisionOffsitenow captures it after the resource is ready via aHostKeyScannerseam (SSHHostKeyScanner, x/crypto/ssh — dials port 23 and grabs the host key from the handshake, no ssh binary needed). Fail-closed: a nil scanner or a scan failure returns an error (don't serve a descriptor the controller can't verify). The controller re-scans and refuses on mismatch (v0.106.0).- Tests: descriptor carries the fingerprint from a faked scanner; a scan failure fails-closed.
v0.37.0 — offsite provisioning SLICE 1: Hetzner Cloud-API client + provisioning core (2026-07-09)
Slice 1 of the offsite-provisioning epic. On operator enable, the hub provisions a Hetzner storage-box
sub-account (shared) or dedicated box, generates the transient password, stores it one-time-consumable, and
serves the non-secret target descriptor to the controller via ConfigJSON. Coded against the API shapes
measured live in documentation/audits/SPIKE-hetzner-api-provisioning-2026-07-09.md (996d403). The
controller apply-bridge (SLICE 2), escrow auto-confirm (SLICE 3), and soft-quota enforcement (SLICE 4) are
separate slices.
internal/hetznerapi: a typed client for the storage-box surface athttps://api.hetzner.com/v1(NOTapi.hetzner.cloud— the classic Cloud API 404s for storage boxes). Sub-account + box create/reset/access/change_type/delete/list-by-label +WaitAction(poll tosuccess, bounded). ACloudAPIinterface + an exportedFakeso provisioning is unit-tested with no live Hetzner calls. Bearer token from an injected func (out-of-band secret; never logged).internal/offsite:Provisioner.ProvisionOffsite— idempotent bylabel_selector(felhom-customer=) (names aren't unique); shared → sub-account on the pool box, dedicated → box; generates a 4-class transient password →WaitAction→Store.SaveOneTimeSecret→ builds the NON-SECRETDescriptor{enabled,type,host,user,port:23,repo_path:/home/felhom-repo, quota_gb|box_type}. Fail-closed: any API/action error returns without a provisioned resource, one-time password, or descriptor.MergeDescriptormerges it under theoffsitekey ofConfigJSON(never a secret).internal/store:one_time_secretstable +SaveOneTimeSecret/ConsumeOneTimeSecret(single-use, return-and-mark in one tx). The transient password NEVER ridesConfigJSON.internal/api:POST /offsite/consume-password/{id}— serves the one-time password to the authenticated customer (same API-key auth as config-pull) EXACTLY once, then 404s. Never logged.internal/web: the config form gains an Offsite backup section (enable / type / soft-quota / box type); save →applyOffsiteprovisions (fail-closed: a provisioning error returns 502 and does NOT save) and merges the descriptor →ConfigVersionbump → controller re-pulls. Optional dep (SetOffsiteProvisioner), wired incmd/hub/main.gofromHETZNER_TOKEN/HETZNER_POOL_BOX_ID/HETZNER_LOCATION.- Tests (faked Cloud API, no live calls): shared/dedicated provision + descriptor + one-time-password-stored
- password-absent-from-ConfigJSON; idempotent re-save (no 2nd resource); fail-closed + companion
red-proof (swallow the create error → offsite marked enabled despite failure → test fails); one-time
consume-once;
WaitActionsuccess/error/timeout; the consume endpoint (auth + single-use).
- password-absent-from-ConfigJSON; idempotent re-save (no 2nd resource); fail-closed + companion
red-proof (swallow the create error → offsite marked enabled despite failure → test fails); one-time
consume-once;
- NOT yet live-provisioned — awaiting the dedicated-project scoped token (the current token can delete ep0 — SPIKE §6); a live create is a supervised validation. Unit tests are this slice's proof.
v0.36.0 — customer page: passphrase hardening + interactive install-command generator (TASK GL-7) (2026-07-09)
Two coupled, security-first changes to the operator-facing customer page (customer_unified.html +
configs.go). felhom.eu only; agent + host-install untouched.
- Passphrase hardening (ships the security win). The per-customer retrieval passphrase was
rendered in cleartext twice — the visible
#retrieval-pwnode and baked into the Option-3 debug curl'sX-Retrieval-Password:header. Now:#retrieval-pwrenders a masked bullet run by default with reveal (toggleSecret) + copy (copySecret) controls, the value carried indata-secret(the existing reveal model). The Option-3 command carries a<YOUR-RETRIEVAL-PASSWORD>placeholder — the secret is NEVER in a copyable command block. (A zero-secret-in-DOM reveal-on-demand fetch is a deliberate future follow-up, not this task.) - Interactive install-command generator. The three hard-coded install
<code>blocks became a client-side builder (vanilla JS — no framework, CDN, or network) that assembles a live-updating command from form controls, emitting ONLY real host-install v1.12.0 flags in a download-then-run shape (nevercurl | bash). CustomerID is prefilled from the server (ScriptVersion/data-customer-idviapageData); a byo selection requires--cores/--memory(enforced client-side with a.gen-req/gen-msgprompt); caps/mode are placeholders, never silent defaults. Graceful JS-off static fallback: the Option-1/2 code nodes keep a--customer-id … --mode <appliance|byo>command. The curated control surface excludes the seven dangerous/operator-only flags (--force,--rotate-recovery,--enable-oob,--remove-golden,--uninstall,--adopt-pool,--rescope-acl) — they are never offered as controls. configs.go:const hostInstallVersion = "1.12.0";pageData.ScriptVersionadded + populated.- Tests (
render_test.go):TestTemplates_PassphraseHardened(secret NOT in the Option-3 command, placeholder present, masked-by-default bullet run,data-secretpopulated, reveal/copy controls present; red-proof = revert Option-3 to the raw secret → fails) andTestTemplates_InstallGenerator(all curated control ids present, script version +data-customer-idrendered, static-fallback command present, and none of the seven excluded flags appear page-wide). - Style (
style.css):.gen-controls/.gen-radios/.gen-radio/.gen-check(s)/number inputs/.gen-msg— dark palette, 2px radius.
v0.35.0 — OOB operator access: operator peer + oob_peer_ip/oob_operator_ssh_key + OOB health alert (TASK H1) (2026-07-05)
The hub half of the merged E1+H1 operator-SSH-access feature (agent half = felhom-agent v0.72.0).
- Operator OOB peer (
store/wg_operator.go): the fleet operator peer as an UNBOUND wg_peers row (host_id '', note operator-oob) at an EXPLICIT /32 (so the endpoint's static forward chain can hardcode it); validated in-subnet/not-reserved/not-taken; last-write-wins rotation. It rides ListWGPeers → peersync pushes it to the endpoint.PUT/GET /admin/wg/operator-peer(global key); the PUT also takes an optionalssh_pubkey(the operator authorized_keys line, hub_settings) and bumps EVERY host's generation. - Desired-state (
api/wg.gomergeWireguard): when an operator peer exists, the served wireguard block carriesoob_peer_ip(rendered into the box's AllowedIPs — survives self-heal [OF-1]) andoob_operator_ssh_key(agent writes felhom-sshd's authorized_keys). Absent → byte-identical. - OOB health (
monitor/host_oob.go): ingests the agent'soobheartbeat stanza and raises a transition-basedoob_degraded/oob_recoveredwarning (felhom-sshd down while the operator peer is configured, OR config invalid) — the proactive "can the operator get in right now" signal.
v0.34.1 — mgmt_plane_healed alerts on the FIRST auto-heal (TASK G1 fix) (2026-07-05)
The mgmt-plane checker seeded a heal marker silently on first observation (copied from HostLeafChecker's trust-on-first-report). A heal is an EVENT, not a baseline: construction still seeds pre-existing markers (startup false-alarm guard), but a newly-observed marker now raises the warning — so the FIRST auto-heal surfaces, matching the live drill. Added tests for both halves.
v0.34.0 — break-glass recovery vault + mgmt_plane surfacing (TASK G1) (2026-07-05)
The hub half of the management-plane break-glass system (prerequisite for felhom-sshd / H1; agent half
= felhom-agent v0.71.0). Closes the recovery gap from
documentation/audits/SPIKE-felhom-sshd-2026-07-05.md §8/#9.
- Break-glass credential vault (
store.host_recovery+internal/store/host_recovery.go): a per-host root@pam console password, stored at rest, operator-retrievable — the human fallback for reaching the PVE web console (pveproxy, a failure domain distinct from sshd) when both the sshd path and the agent-independent auto-heal have failed.PUT /hosts/{id}/recovery-credential(SELF-scoped host key — day-0 vaults it) +GET /admin/hosts/{id}/recovery-credential(GLOBAL key only — a host key cannot read its own console password back). Secret discipline: never logged (username + length only); red-proofed that the password never reaches the hub log. - mgmt_plane surfacing (
internal/monitor/host_mgmtplane.go, on the 60s sweep): parses the agent's additivemgmt_planeheartbeat stanza and raises amgmt_plane_healedWARNING when the watchdog auto-healed a missing/run/sshd(newprivsep_healed_at) — a recurring clobber surfaces BEFORE it becomes a lockout, complementing host_staleness. Trust-on-first-report (seed, then alert on change), mirroring HostLeafChecker.
v0.33.0 — S2 offsite connectivity: box-facing WG registration + wireguard desired-state block + /offsite UI (2026-07-04)
Doc 06 roadmap row S2 (commits fcf84a0/ba52005/13203c2); the S2 architectural decision:
the stored desired_json stays a pure OPERATOR blob — the WG assignment is HUB-owned state,
merged into the served desired-state at READ time, never written into the store.
- Store (
internal/store/wg.go+store.go):RegisterWGPeerForHost(idempotent / re-key-in-place keeps the /32 — stable tunnel addressing across rotation/DR / adopt-unbound S1 rows / typedErrWGPubkeyBoundElsewhere— a key is never silently stolen); partial unique indexidx_wg_peers_host= one bound peer per host;BumpHostDesiredbumps ONLY the generation (the merge changes served state, not the blob);allocateWGPeerTxextracted from the S1 path behavior-neutrally (S1 tests unmodified).WGPeergainsCreatedAt. - API (
internal/api/wg.go+handler.go):POST /hosts/{id}/wg— per-host key SELF-SCOPED (global = operator/DR path); generation bump + endpoint push ONLY on real change (idempotent re-register moves nothing — asserted negatives).mergeWireguardinjects{endpoint{dns_name,wg_port,server_pubkey,pbs_tunnel_ip}, pubkey, assigned_ip}into served desired-state; no peer → byte-identical pass-through (the cross-repo golden test passes UNMODIFIED); any merge failure → fail-safe unmerged serve (never 500 the control channel).handleAdminSetDesiredStateREJECTS a top-levelwireguardkey (400 — an operator copy-paste-PUT can never clobber the hub-owned block). Admin DELETE of a BOUND peer bumps the owning host; unbound deletes move no generation. NEW goldentestdata/desired-state-wireguard.golden.json= the S3 cross-repo contract (agent copy must stay byte-identical). - UI (
internal/web/offsite.go+templates/offsite.html): read-only/offsitepage — endpoint card + peer table (truncated pubkeys, full value in title; bound peers link to/hosts/<id>); Offsite nav link in all 9 page templates. Mutations stay on the admin API (UI actions arrive with tunnel health, S3/S6). - Tests: Groups A/B/C; five red-proofs run + reverted (self-scope drop, unconditional merge, rejection drop, bump-on-idempotent, script exit-swallow — see scripts/CHANGELOG v1.0.1). Old-agent (v0.63.0) tolerance proven live against the real felhom-pve record.
v0.32.0 + v0.32.1 — S1 offsite connectivity: WG endpoint record + peer registry + pinned-SSH peer-sync (2026-07-04)
The hub side of doc 06's roadmap row S1 (documentation/architecture/06-offsite-connectivity.md),
resolving the slice-1 design point: peer-sync = hub pushes over SSH to a forced-command
reconcile script on the endpoint (pull/signed-manifest rejected — weakens immediate revocation;
HTTPS push API rejected — a new versioned binary + third public port for nothing).
- Store (
internal/store/wg.go+ migration instore.go, commitb18f6ae):wg_endpoints(single expected row "ep0") +wg_peers(presence = desired state; no status column — that's the S2 host-join).AddWGPeer= one tx, idempotent on pubkey, lowest-free-host/32allocation skipping network/pbs_tunnel_ip/broadcast,UNIQUE(assigned_ip)race backstop + one internal retry; typedErrWGEndpointUnset/ErrWGSubnetExhausted. - wgsync (
internal/wgsync/, commitsfbeeacb+0fa7ea1):x/crypto/sshpush client withssh.FixedHostKeypin (no insecure fallback, ever) +HostKeyAlgorithmsconstrained to the pinned key's type — the live validation caught a stock multi-hostkey sshd presenting ECDSA against the ed25519 pin (legitimate server refused); regression-tested with an in-process dual-hostkey SSH server. Reconciler pushes the FULL peer list (never deltas — drift repair by construction) onTrigger()or a 5-min tick; payload{"version":1,"interface":"wg0","peers":[{pubkey, allowed_ip}]}, deterministic order. - API (
internal/api/wg.go):PUT/GET /admin/wg/endpoint,POST/DELETE/GET /admin/wg/peers— GLOBAL key only (thehandleAdminSetDesiredStategate); pubkey validated 44-b64/32-byte; DELETE takes the pubkey in the JSON body (base64/++keep pubkeys out of URL paths); mutation responses carrysync: ok | deferred: <err> | disabled— the DB is the source of truth, a failed push defers to the reconciler. - Wiring (
cmd/hub/main.go):WG_ENDPOINT_SSH_{ADDR,USER,KEY_FILE,HOSTKEY}env (key from the mountedSecret/wg-endpoint-ssh, host key non-secret plain env); any piece missing →[INFO] WG peer-sync disabledand mutations still work DB-only. - Tests: allocator (exact IPs, freed-IP reuse, /30 exhaustion), API auth/validation with a fake syncer, SSH client against an in-process server (exact payload bytes, stderr surfacing, wrong-host-key refusal, multi-hostkey pin), reconciler (full-list, retry-on-tick, no-mutation drift push, removed-peer-absent negative). Four red-proofs run and reverted (allocator-ignores- rows, gate removal, InsecureIgnoreHostKey, delta-only push) — each failed its test.
- Live-validated end-to-end on the dev endpoint (
felhom-hetzner, runbookdocumentation/runbooks/offsite-endpoint.md): add →wg showon the box; delete → gone (+404/403 paths); malformed payloads leave wg state byte-identical; endpoint reboot → persisted set + hub push converges; client tunnelep0.felhom.eu:443→ PBS login page via the wg0-only 8007 rule; public 8007 unreachable. v0.32.1 = the HostKeyAlgorithms fix (0.32.0 image was already pulled by the cluster; tag kept immutable).
docs — Felhom skills introduced + CLAUDE.md refresh (2026-07-03)
Repo-level docs work alongside v0.31.0 (no hub code in this entry):
skills/(new, repo root): three versioned Claude Code skills —felhom-build-deploy(per-artifact runbooks, all commands verified live),felhom-ui-design(v2 tokens + gates),felhom-testing(non-hollow doctrine + red-proof procedure). Installed to~/.claude/skills/viascripts/install_skills.py(junction mode verified).- CLAUDE.md refresh: the "Hub — current state (v0.7.x)" narrative (stale by ~23 versions) replaced with a version-free architecture section; standing rule adopted — CLAUDE.md carries NO version-pinned state (that lives in CONTEXT/CHANGELOG/REUSE); skills pointers added. Same rule applied to the sibling repos' CLAUDE.md in their own commits.
v0.31.0 — critical severity accepted at event ingest + visible in UI (2026-07-03)
Fixes the gotcha the REUSE sweep surfaced: handleEvent coerced any severity outside
{info,warning,error} — including "critical" — to "info" at ingest, so a controller-POSTed
critical event never notified even though the dispatcher (severityNotifies, v0.24.0) and
FormatOperatorEmail already handle critical correctly.
- Ingest (
internal/api/handler.gohandleEvent):"critical"added to the severity case list. Unknown values (and case-variants like"Critical") still coerce to"info"— the exact-match-lowercase coercion contract is kept and now locked by test. - Hungarian label (
internal/notify/templates.go):severityLabels["critical"] = "Kritikus hiba"(was missing — customer emails would have shown the raw English word). - UI counts: dashboard consumer (
internal/web/server.go) gainsEventCriticals;dashboard.htmlrenders the critical badge FIRST in the 24h count chain (guard extended);customer_unified.htmlgains the{{.}} criticalsummary badge before errors. - style.css: defines the previously-referenced-but-undefined
.severity-critical(--crit/--crit-dimtokens) and.severity-ok(neutral, exception-color principle). No other restyle. - Tests (
internal/api/event_test.go, new): critical preserved to store (companion red-proof: shown failing against the pre-fix switch — stored"info"); unknown severity → info; unknown event_type → 400 + nothing stored; no-auth → 401. First tests on the /event endpoint. - REUSE.md §1/§3 updated in the same commit (the maintenance rule's first outing).
docs — REUSE.md introduced (2026-07-03)
Cross-repo reuse-map rollout (docs-only, no code change, no version bump). New REUSE.md at the
repo root covering hub + website + scripts + manifests: canonical helpers (34 rows), patterns
(monitor checker, website page, gate script, GitOps deploy), dangerous lookalikes (legacy /notify
trio, severity-critical coercion at handleEvent ingest, inline stringData secrets, kubectl-apply
drift…), seams, extension points, and observed duplication (5 clusters, NOT fixed). New
scripts/reuse_refs_check.py machine-checks every cited path in all four repos' REUSE.md files.
CLAUDE.md gains the REUSE.md pointer + same-commit maintenance rule.
v0.30.1 — status badge no-wrap (2026-07-02)
Found in the authenticated D4 validation pass: multi-word status tags (PENDING in a narrow
dashboard column, NO REPORT on hosts) wrapped between the CSS dot and the label. One line:
white-space: nowrap on .status-badge.
v0.30.0 — TASK-D4: design system v2 re-skin (appearance only) (2026-07-02)
Last surface of the design sprint (controller D0/D1, website D3). The hub leaves its Tailwind-slate
theme for the canonical navy v2 language. API surface untouched (/api/* ingestion, artifact
manifest, config generation, DR/escrow — git diff clean under internal/api + internal/store).
- Fonts (
internal/web/static/fonts/, embed.go, server.go): the 4 vendored woff2 (byte-copied from felhom-controller; latin-ext for Hungarian customer names in an English UI), embedded and served at/static/fonts/(font/woff2, immutable), mirroring the chart.min.js pattern. No CDN before or after. statusColorsemantic remap (server.go): returnsnominal/warn/crit/neutralclass tokens instead of raw hex colors — ok→nominal, warn+stale→warn, down+fail→crit, pending+disabled→neutral (a not-yet-provisioned or deliberately paused customer is a normal fleet state), blocked→warn (intentional operator cut-off: attention-worthy, not an outage). The inlinestyle="color: {{statusColor}}"pattern is dead (dashboard + customer_unified use class-based.status-dot-<token>);statusIcon("●") retired. Truth-table test red-proven vs the old implementation; new template-parse test (neither existed for the hub).- style.css v2: navy tokens + @font-face; 2px radius; hairline
--line-softtable rows (fleet-NOC density kept);.status-badgere-expressed as an outline tag + CSS dot per the design-system addendum (ok=blue, warn/blocked/stale=amber+dim, down/fail=red+dim, pending/disabled=quiet neutral with hollow dot); severity badges stay filled amber/red (exceptions stay loud); config badges = filled informational chips in v2; row tint only for warn/down. Two-tone brand H1 (Felhom <span>Hub</span>) on all pages; 12-symbol Lucide sprite partial included per page. - Charts (app_detail): avg memory
#2EA8F5, peak#8E7CE8(secondary DATA series — not status red), catalog-limit line#E0A93E(threshold marker); legend/tick/grid → v2 literals. - customer_unified JS status-message colors → blue-bright/crit; login page inline HTML retinted.
- Grep gate: all slate hexes (
#0f172a #1e293b #334155 #60a5fa #4ade80 #facc15 #f87171 #94a3b8 #64748b #475569 #e2e8f0) at zero across internal/web (non-test).
v0.29.0 — Day-0 artifact manifest: version dropdowns + auto-derived sha (2026-07-01)
Removes the hand-copied sha256 from the Day-0 artifact manifest. The operator now picks a version from a dropdown of what's actually in Gitea (olders get pruned), and the hub reads that version's sha256 from Gitea itself — no transcription, no stale checksums. Keeps the human-in-the-loop trust gate (the operator still deliberately chooses the version; "latest" is never auto-promoted) while the hub stays the checksum trust root.
internal/gitea(new): a minimal read-only Gitea packages client —ListVersions(generic package versions, newest-semver first) +FileSHA256(a version's file sha256 via the files-metadata API, without downloading the artifact — important for the ~GB golden). Basic-auth with the registry creds the hub already holds. Unit-tested against an httptest server (filter+sort, preferred file match + fallback, non-200 → error).- Configuration → Day-0 artifacts: the two version text inputs are now
<select>dropdowns populated from Gitea; the sha256 fields are read-only, displayed (mirrored from the picked version via a tiny inline script). Choosing "— none —" clears an artifact. handleSetArtifacts: derives each chosen version's sha256 from Gitea authoritatively (a client-submitted sha is ignored); a Gitea lookup failure REFUSES the save (never stores a version with a wrong/blank checksum) rather than corrupting the manifest.- Graceful degradation: with no registry creds (
web.SetGiteaClientnot wired) the form falls back to the previous manual text-entry path.main.goenables the Gitea browser whenREGISTRY_USERNAME/REGISTRY_TOKENare set. go build/vet/test ./...clean.
v0.28.0 — global settings → Configuration tab + online setup command (2026-07-01)
Three operator-requested improvements (companion: host-install script v1.2.0).
- Global settings moved from the Customers page to the Configuration tab (
web/configuration.html,configs.html,server.go,configs.go). The two global cards — "Managed updates — global floor" and "Day-0 artifacts — agent & golden" — were on the Customers list; they now render + save on Configuration (where they belong).handleConfigurationsuppliesGlobalFloor+Artifacts+CSRFField; the save handlers (handleSetGlobalFloor/handleSetArtifacts) now redirect to/configuration?flash=…and are mounted at/configuration/global-floor+/configuration/artifacts; their flash banners moved too. The Customers page is back to just the list + "Add Customer" (its per-customer effective-floor column is unchanged). - Online setup command added to a customer's Setup Command block (
customer_unified.html). New "Option 1: Online install (recommended)" — download-then-run:curl -fsSL https://felhom.eu/scripts/felhom-host-install.sh -o … && sudo bash … --customer-id <id>— with a copy button and the customer id filled. The passphrase is not templated in (entered at the prompt). The former local-file command becomes Option 2, the debug curl Option 3. Download-then-run (notcurl | sudo bash) stays the recommended form — inspect before running. - Website now serves
/scripts/(manifests/webpage.yaml). The host-install script lives at the repo's/scripts(outside the website doc-root); added/scripts/to the git-sync sparse-checkout and an nginxlocation /scripts/(root.../current,text/plain) sohttps://felhom.eu/scripts/felhom-host-install.shresolves — single source of truth, no duplicated copy. - Remaining audit follow-ups unchanged: controller-side geo intent sync; a read-only reported-vs-desired
"Show Diff"; the cosmetic
controllerURLcleanup inconfigs.go.
v0.27.0 — Hosts page: read-only fleet view (audit F-M1) (2026-07-01)
Resolves audit finding F-M1: the agent enrolls as a host and the hub stores rich host state (identity, agent version, guests, storage targets with SMART, DR/escrow, staleness) and alerts on it — but the whole host domain was invisible in the GUI (email-only). Adds a Hosts nav section — a fleet list + a per-host detail page. Read-only (GET only, no host actions/mutation routes): this surfaces state the way the pull/desired-state model demands; it does not reintroduce inbound control (retired in v0.26.0).
- New store reader
ListGuestsForHost(hostID)(internal/store/store.go).SELECT … FROM guests WHERE host_id = ? ORDER BY vmid, via a newscanGuesthelper overguestRealitySelectCols— the reality columns only. It deliberately omits the secret/inert columns (api_key,desired_spec_json), so the read-only view can never surface them. Returns[](never nil-error) on no guests. (There was previously no guests reader — onlyUpsertGuestFromReport.) - New handlers (
internal/web/hosts.go).handleHostsList—ListHosts+ a per-host status badge fromhostStatus()(which reusess.staleThreshold— the same thresholds as theHostStalenessChecker: stale after the threshold, down at 2× — so the badge agrees with the alerting) + per-host guest counts (ListGuestsForHost) + vitals parsed fromGetLatestHostReportJSON+ worst storage fill grouped fromGetHostStorageTargets.handleHostDetail—GetHost(404 if absent) +ListGuestsForHost+ rich storage targets (role, state, fill %, thin-pool, SMART health/temp/wear parsed from the latest report body) + vitals +GetHostDRBundle/GetHostEscrowpresence booleans only (never the opaque blobs). Nil/missing (no report, no guests, no storage, no DR) render empty states — never a panic.
- New templates
hosts.html+host_detail.html(existing dark operator-console styling reused —data-table,status-badge-*,info-grid,empty-state; no restyle). A no-report host shows a STALE/NO-REPORT badge and "waiting for first report". - Nav: added the
Hostslink (between Apps and Configuration) to every page's<nav>(the nav is duplicated per page, not a shared partial) + the newtimeAgoPtrtemplate helper for*time.Time. - Routes (
internal/web/server.go):GET /hosts(+/hosts/) → list,GET /hosts/{id}→ detail, modelled on the/appspair. GET only. - Tests:
ListGuestsForHost(none→empty, multiple→vmid-ordered, secret column not surfaced); the list handler (N rows, ONLINE + NO-REPORT badges, worst-fill, no action buttons); the detail handler (guests + storage + SMART + DR present, customer cross-link, no-secret assertion that the hostapi_keyis absent from the rendered body, no buttons); unknown host → 404; no-report host renders the waiting state.hostStatusband mapping unit-tested (pending/ok/stale/down). - Remaining audit follow-ups (not this slice): controller-side geo intent sync; a read-only
reported-vs-desired "Show Diff"; the cosmetic
controllerURLcleanup inconfigs.go.
v0.26.0 — pull-based config delivery + retire the inbound GUI controls (2026-06-30)
Closes audit documentation/audits/AUDIT-hub-gui-2026-06-30.md F-S1/F-S4 + the dead-template findings,
and replaces the never-inbound-violating "Push Config" with a pull-based config-refresh that rides the
report ACK (companion controller change: felhom-controller v0.94.0).
- Config delivery is now pull-based (
internal/store/store.go,internal/api/handler.go). Newcustomer_configs.config_versioncolumn — a stored counter (NOT a hash of the rendered YAML;configgenemits a freshweb.session_secret+ timestamp every call, so a content hash would change spuriously).SaveCustomerConfigbumps it on every save (new rows seed at 1, updates increment) — the one path that changes the generatedcontroller.yaml(identity + theconfig_jsonoverrides). The floor, block/unblock, and retrieval-password regen deliberately do NOT bump it. The report ACK (handleReport) now advertisesconfig_versionbesidemin_controller_version/latest_version; the controller compares it to its last-applied version and re-pulls + self-restarts on a change. Omitted for report-only (no-config) customers, so an old controller is unaffected. - Retired the five inbound (hub→box) controls that violated the never-inbound posture
(
01-topology-and-trust.md:11) and were broken behind the box's CF tunnel/NAT:- Trigger Update — handler + route deleted; controller updates are agent-driven (the version floor).
- Push Config — handler + route deleted; replaced by the pull-based config-refresh above.
- Pull Config — handler + route deleted.
- Show Diff (
handleConfigDiff+ thecompareYAMLValues/flattenYAML/maskSensitivehelpers) — deleted, along with the now-deadConfigSyncStatus/ConfigDiffCountplumbing. - Geo-disable — KEEPS its legitimate hub→Cloudflare WAF-rule removal (
RemoveGeoRules); the secondary inboundnotifyControllerGeoDisableis deleted. After this,grep client.Do internal/web/has zero ControllerURL targets (only Gitea registry/template fetches remain; the ControllerURL is still shown as a display-only link).
- GUI staleness (F-S1) + dead templates: the customer page's Setup Commands now show the Proxmox
Day-0 host bootstrap (
sudo ./felhom-host-install.sh --customer-id <id>, passphrase at the no-echo prompt) instead of the pre-Proxmoxdocker-setup.sh; Option 2 relabelled "Manual config fetch (debug only)". Deleted the orphanedcustomer.html+config_detail.html(rendered by nothing;/configs/{id}redirects to/customers/{id}). - Audit doc (deferred line): the GUI audit
documentation/audits/AUDIT-hub-gui-2026-06-30.md(committede51e03b) is the grounding for the above; its F-S1/F-S4 + dead-template findings are now resolved. Open follow-ups noted there remain: the Hosts page (F-M1), controller-side geo intent sync, and Show-Diff could return later as a read-only-vs-reported view. - Tests: store
config_versionbump (create=1, edits increment, per-customer independent) + the no-bump red-proof; ACK carriesconfig_versionand omits it for report-only customers + the no-bump red-proof.go build/vet/test ./...green.
v0.25.0 — per-storage worst-fill alerting (StorageFillChecker) (2026-06-30)
Generalizes the host-root disk alert (v0.23.0) to ANY reported storage target — so a dedicated dump/backup volume, data drive, lvmthin pool, or PBS datastore filling toward failure pages the operator with the storage named, even when host root itself is fine.
internal/monitor/storage_fill.go(NEW) —StorageFillChecker. A per-target mirror ofHostDiskCheckeron the same 60s sweep: born/persistent (already-breached(host,target)keys left UNSEEDED → firstCheckemits), escalation-only emit, recovery re-arm, the dispatcher's 1h cooldown. State is keyed per (host, target) so targets alert independently. Emits distinctstorage_fill_warning/storage_fill_criticalat the naturalcriticalseverity (exercises the v0.24.0 dispatcher fix with a second real caller). Default thresholds 90/95, hub-config overridable (alerting.storage_fill_warn_percent/_crit_percent), independent of the host-root thresholds.- Root excluded (no double-alert): the host root-backed builtin (
Type=="local", or a target mounted at/) is skipped —HostDiskCheckerowns root. So a root-backed vzdump dump is one alert (from host_disk), and storage_fill uniquely covers OFF-root storage. internal/store/store.go:GetHostStorageTargets()+HostStorageTargetRow— parsesreport_json.storage_targets[]of each host's latest report (percent =used_fraction×100); modeled onGetHostDiskUsage, no denorm column / migration.internal/notify/templates.go+internal/api/handler.go: Hungarian templates + allowlist entries forstorage_fill_warning/storage_fill_critical.cmd/hub/main.go: registerstorageFillCheckeron the 60s tick besideHostDiskChecker.- Tests: per-target bands (independent warn/escalate/recover/re-arm), born/persistent companion red-proof
(a seed-all model stays silent on the born-breach), root-exclusion companion (without the exclusion a
root target IS in the critical band — the exclusion is what suppresses the double-alert), severity
critical, and the store parse.go build/vet/test ./...green.
v0.24.0 — dispatcher routes critical severity (+ nil-prefs crash guard) (2026-06-30)
NAS Part A2's "Part 0": close the dispatcher's silent drop of critical-severity events.
internal/notify/dispatcher.goProcessEvent: the severity gate wasseverity != "warning" && severity != "error"→ acriticalevent was silently dropped (never emailed). Now routes warning / error / critical (severityNotifies);infostays an intentional non-notify; any unrecognized severity is logged ([WARN] Dispatcher: unrecognized severity …), never silently dropped. Verified safe first: no controller event emitscritical(all are info/warning/error) and the hub's only would-becriticalemitter ishost_disk— so no surprise alert volume.internal/monitor/host_disk.go:host_disk_criticalnow emits its naturalcriticalseverity (was forced toerrorto survive the old gate);FormatOperatorEmailstylescritical🔴 likeerror.- Latent crash guard:
processCustomerdereferencedGetNotificationPrefs, which returns(nil, nil)for a customer with no notification row — an event for such a customer would panic the dispatcher goroutine and crash the hub. Now guardsprefs == nilbefore use. - Seam:
sendEmailFnfield (defaults to the Resend sender) so routing is unit-tested without real HTTP. - Tests:
severityNotifies(warning/error/critical notify; info/unknown don't) + companion red-proof (the pre-fixwarning||errorpredicate dropscritical); ProcessEvent routescriticalto the operator; an unknown severity is logged not dropped;infois silent and not mis-logged.go build/vet/test ./...green.
v0.23.0 — host root-disk pressure monitoring + alert (2026-06-30)
Closes the silent-failure gap behind the felhom-pve incident: a Proxmox host root fs filling up (vzdump
piling under /var/lib/vz/dump) went unnoticed because nothing alerted on the HOST root disk_percent the
agent already reports. New hub-side checker on the existing 60s sweep.
internal/monitor/host_disk.go(NEW) —HostDiskChecker. A sibling ofHostCapabilityChecker/HostLeafChecker: reads each host's latest root-fsdisk_percent(store.GetHostDiskUsage) and emits an operator alert on a warning (default 90%) or critical (default 95%) crossing. Rank-based bands (ok→warning→critical) so an escalation always alerts and a de-escalation/recovery re-arms silently.- Born/persistent (the F2 lesson): a disk ALREADY over threshold when the hub/checker (re)starts
alerts on cycle 1 — seeding leaves already-breached hosts UNSEEDED so the first
Checkemits (a transition-only design would stay silent forever on a persistently-full disk). The dispatcher's 1h operator cooldown dedups re-emits across a hub restart. - Distinct event types
host_disk_warning/host_disk_critical— NOT the controller's GUESTdisk_warning/disk_critical(the guest cgroup view), so the host and guest alerts never dedup or mask each other. - Severity: warning band →
warning; critical band →error(NOT"critical"). The dispatcher only routeswarning/errorseverities — a"critical"severity would be silently dropped — so the critical band maps toerror(and the operator email's 🔴). (Deviation from the task's stated "critical → critical", made to match the live dispatcher.) - Thresholds are hub-config overridable (
alerting.host_disk_warn_percent/host_disk_crit_percent, seed-only); an unset/invalid/misordered config falls back to 90/95 (normalizeDiskThresholds) so a typo can never silence or invert the alert.
- Born/persistent (the F2 lesson): a disk ALREADY over threshold when the hub/checker (re)starts
alerts on cycle 1 — seeding leaves already-breached hosts UNSEEDED so the first
internal/store/store.go:GetHostDiskUsage()+HostDiskRow— latest report per host (MAX(id)),disk_percentfrom the denorm column + total/used bytes parsed fromreport_json(event detail). No schema migration.internal/notify/templates.go: Hungarian customer templates forhost_disk_warning/_critical(customer delivery still requires per-customer opt-in via enabled events; operator alert is the headline).internal/api/handler.go:host_disk_warning/host_disk_criticaladded toallowedEventTypes.cmd/hub/main.go: registerhostDiskCheckeron the shared 60s tick.- Tests: band transitions (seed/escalate/steady/recover/re-arm), severity mapping, threshold defaults, and
the born/persistent companion red-proof (a seed-all/transition-only model stays silent on a
born-breach; the real unseeded design emits).
go build/vet/test ./...green. - Follow-ups (noted, not built): per-storage
StorageTargetsworst-fill alerting (a dedicated dump/backup storage filling — host rootdisk_percentalready covers the observed case); and the provisioning-sideprune-backupsretention default so a box can't refill its own root (operational fix, separate from this detector).
hub-config — enable operator email alerts (config-only, no image change) (2026-06-30)
manifests/hub.yaml hub-config ConfigMap: set notifications.operator_email: admin@felhom.eu +
operator_enabled: true. The dispatcher's operator path (Dispatcher.processOperator) sends only when
operatorOn && operatorEmail != ""; without these it returned early, so the self-health pipeline
(probe → report → checker → dispatch) stopped one hop short of the inbox (the TESTRUN's unproven hop).
No hub image change (live tag stays v0.22.1) — ConfigMap edit + pod restart to reload.
- Proven end-to-end (2026-06-30): a real capability-degrade alert produced
[INFO] Operator email sent for demo-felhom/agent_capability_degraded(the send-success line that never fired while the path was gated); the customer path (POST /api/v1/notifyevent_type:test) sent to the customer address via the samesendEmail→ Resend. Seedocumentation/audits/TESTRUN-fullstack-2026-06-29.md("Findings closed", Part A). Operator email is not a secret; the Resend key stays injected fromSecret/resend-api.
v0.22.1 — wire HostLeafChecker into the monitor loop (v0.22.0 missed the wiring) (2026-06-29)
The v0.22.0 commit added HostLeafChecker but the cmd/hub/main.go goroutine edit never applied, so
the checker was never started. The live test caught it (a leaf regen produced no host_leaf_changed,
only the controller's complementary agent_channel_pin_mismatch). Now started on the 60s sweep next to
the staleness/capability checkers. No other change.
v0.22.0 — proactive agent re-key detection: HostLeafChecker (host_leaf_changed) (2026-06-29)
Companion to felhom-agent v0.48.0 (which now reports its served local-API leaf fp). The hub watches each host's leaf fingerprint and raises an operator alert when it changes (an agent re-key) — proactive, fleet-wide, independent of any controller's channel-health check. The last self-health leg.
monitor.HostLeafChecker(NEW): sibling ofHostCapabilityChecker. Trust-on-first-report — the first fp per host is the baseline; a later change emitshost_leaf_changed(operator, English; details carry old+new fp) and advances the baseline. First-obs seeds silently (a "change" needs a prior value, so no F2 issue). An empty reported fp (pre-v0.48.0 / local-API-disabled) is unknown — never seeds, never alerts, never overwrites a baseline. Customer-blocked hosts dropped; unseen pruned. Runs on the existing 60s sweep. KNOWN LIMITATION (documented): trust-on-first-report can't detect a re-key that happened before the hub's first report — but the controller channel-check catches the downstream pin mismatch, so this is defense-in-depth.store.GetHostLeafFingerprints(NEW): latest reported fp per host, parsed fromreport_json(mirrorsGetHostCapabilities—MAX(id), no schema migration).- No allowlist change:
host_leaf_changedis hub-GENERATED (viaSaveEvent+dispatcher.ProcessEvent), not controller-pushed, so it bypasses the/api/v1/eventallowedEventTypesgate — same as the host_* events. The generic operator template relays it (no template change). - Tests: change red-proof (A→B → one event + baseline advanced; companion: unchanged → none),
first-obs seeds silently, change-back re-alerts, empty fp skipped, customer-blocked dropped. Cross-repo
golden mirrors
leaf_fingerprint. Version0.21.0 → 0.22.0.
v0.21.0 — F2: alert on a host already degraded/stale at hub (re)start (2026-06-29)
Mirror of the controller's F2 fix, for the hub checkers: a host that was already degraded
(capability) or stale/down (staleness) when the hub (re)started was seeded silently and never
alerted. Now the constructors seed only HEALTHY hosts; an unhealthy host is left unseeded so the
first Check() emits once. The dispatcher's 1 h operator cooldown dedups the re-emit across a hub
restart (so a hub bounce doesn't re-page for an already-known issue within the window).
internal/monitor/host_capability.go+host_staleness.go: constructor seeds only the healthy state;Check()first-obs (oldState=="") emitsemitTransition(…, "unknown", newState, …)when the observed state isn't healthy, else seeds silently.- Tests: born-degraded red-proof (host degraded at construction → unseeded → one
agent_capability_degradedon first Check, no duplicate on the next); the staleness test updated to the F2 behavior (born-stale → unseeded → onehost_staleon first Check). Version0.20.0 → 0.21.0.
v0.20.0 — Accept controller agent_channel_* events (channel-health relay) (2026-06-29)
Companion to felhom-controller v0.90.0's controller→agent channel health-check. The controller pushes
classified agent_channel_* events to /api/v1/event, but the handler's allowedEventTypes
allowlist rejected them (HTTP 400). Added the 8 types (pin_mismatch, unauthorized, unreachable,
timeout, misconfigured, construction_error, unknown, recovered) so the operator relay works.
These are operator-only (not customer notification toggles, same as the host_* events); the
existing dispatcher routes them generically (no template change). Version 0.19.0 → 0.20.0.
v0.19.0 — Agent capability-degraded operator alert (HostCapabilityChecker) (2026-06-29)
Companion to felhom-agent v0.44.0's privileged-capability self-probe: the agent now rides a
capabilities snapshot on its host report (each required sudo -n grant: ok/degraded), and the hub
alerts the operator when a host transitions into a degraded state — closing the loop that let five
non-root-cutover regressions go undetected until user-visible breakage.
monitor.HostCapabilityChecker(NEW): a deliberate SIBLING ofHostStalenessChecker— same per-host state map (ok/degraded), seed-without-event, emit-only-on-transition shape. A host isdegradediff its latest report has any Critical capability withstatus:"degraded"; non-critical degradations ride the report but never alert. Runs on the existing 60s sweep next to the staleness checkers.- Events:
agent_capability_degraded(warning) on ok→degraded, naming the degraded capabilities- gated features in the message and details JSON;
agent_capability_recovered(info) on degraded→ok. Routed through the existingDispatcher.ProcessEvent— operator-only (the type is not a customer notification toggle, same ashost_stale) with the standard 1 h operator cooldown (no per-cycle re-alert).
- gated features in the message and details JSON;
store.GetHostCapabilities(NEW): reads the capability snapshot from the latest host-report'sreport_jsonper host (keyed onMAX(id)— within-secondreceived_atties would otherwise return multiple rows). No schema migration — the array rides the existing report body. A pre-v0.44.0 agent (nocapabilities) reads asok, so an old agent can't trip a false alert.- Cross-repo
host-report.golden.jsonmirrors the newcapabilities: []field (byte-identical with the agent copy). Version0.18.0 → 0.19.0.
v0.18.0 — App-email passthrough: POST /api/v1/mail → Resend SMTP (2026-06-29)
The hub can now relay a customer box's outbound app email to Resend, re-emitting the raw MIME
unchanged over SMTP. This is the hub leg of the app-email relay (apps → on-box shim → hub →
Resend); the Resend key stays hub-side. Implements documentation/audits/SPIKE-smtp-app-relay-2026-06-28.md.
- New
internal/mailrelay/relay.go:ResendSMTP(aSender) — STARTTLS tosmtp.resend.com:587,AUTH LOGIN resend/<key>(a smallnet/smtp.AuthLOGIN impl; stdlib ships PlainAuth only), then rawMAIL/RCPT/DATA. Raw passthrough — NOT theinternal/notifyResend HTTP-API path, which is unchanged for the hub's own structured alerts and silently drops inline CID images (spike §4). No new external dependency (stdlibnet/smtp). - New
POST /api/v1/mail(internal/api/mail.go): authenticates the box (checkAuthCustomer), enforces the From-header domain allowlist (backstop; reject 403), applies a per-customer in-memory token-bucket rate limit (default 30/min → 429 so one box can't drain the shared Resend quota), then passes the raw bytes through to Resend. Success→200, Resend failure→502 (the box's shim maps that to the app). - Config: new
mailsection (per_customer_per_minute,from_domains); wired incmd/hub/main.goonly when a Resend key is present (else the endpoint returns 503). The existingnotify/dispatcher.goalert path is untouched. - Tests: passthrough byte-equality (raw bytes reach the sender unchanged, not parsed), From-reject + companion red-proof, per-customer rate-limit + isolation + companion, send-failure→502, 401/503/400 paths, token-bucket unit (injected clock), LOGIN auth + From-domain parsing.
v0.17.0 — Resend key sourced from a Secret, out of git (2026-06-29)
Resend rotation + de-git hygiene. The hub's Resend API key was committed in plaintext in
manifests/hub.yaml's hub-config ConfigMap; it is now sourced from an out-of-band Kubernetes
Secret (resend-api) and never lives in git.
cmd/hub/main.go: newRESEND_API_KEYenv override fornotifications.resend_api_key, mirroring the existingREGISTRY_TOKEN/DEFAULT_MIN_CONTROLLER_VERSIONk8s-Secret override pattern. When set, it wins over the (now empty) ConfigMap field. No behaviour change when unset.- The committed ConfigMap
resend_api_keyis now an empty placeholder pointing atdocumentation/runbooks/secrets.md; the live value is injected fromSecret/resend-apivia env. - Part of the Resend key rotation: the previously-exposed send-scoped key was rotated; the new key lives
only in the out-of-band store and
Secret/resend-api(created imperatively, never committed). See the secrets runbook. No secret value appears in this repo.
v0.16.0 — Day-0 artifact manifest (hub-vouched agent + golden checksums) (2026-06-28)
The hub now serves a passphrase-authed artifact manifest so the host-bootstrap script can fetch-then-verify the agent binary + golden archive from Gitea before installing them. The hub is the checksum trust root — a different root than Gitea (which only stores the bytes). Part of the BUNDLE slice that lets a fresh PVE box self-install the agent (no more manual binary/unit step).
store.go:ArtifactManifest{agent_version, agent_sha256, golden_version, golden_sha256}withGet/SetArtifactManifest, persisted as four discrete rows in the existinghub_settingskey/value table (no schema change — same mechanism as the controller-version floor; survives restarts; partial sets round-trip). Added genericgetSetting/setSettinghelpers.handler.go: newGET /api/v1/artifacts/{customer_id}— auth mirrorshandleConfigRetrieveEXACTLY (X-Retrieval-Password, 404-then-401 order, constant-time compare). Returns{"agent":{version,sha256},"golden":{version,sha256}}. An unset manifest returns 200 with empty fields (not an error) so the script falls back to the local golden / fails clearly on a missing binary. v0.16.0 returns the GLOBAL current set for every customer (per-customer pinning is a future hook).configs.go+configs.html(operator UI): a "Day-0 artifacts — agent & golden" card beside the managed-update floor controls, with version + sha256 fields for each artifact.POST /configs/artifacts(CSRF-protected); versions validated as bare semver (reusing the floor validator), sha256s as 64-hex, blank-to-clear.main.go(env seed): seeds the manifest fromARTIFACT_AGENT_VERSION/ARTIFACT_AGENT_SHA256/ARTIFACT_GOLDEN_VERSION/ARTIFACT_GOLDEN_SHA256on startup — but only fields the DB doesn't already have, so a UI edit sticks across restarts. Same escape hatch the Phase-2 floor uses (the operator UI is password-gated).- Tests (
artifact_test.go): recorded set returned verbatim; unset → 200 empty; wrong/missing passphrase → 401; unknown customer → 404; store partial-set round-trip. Plus the floor-render smoke test updated for the new list-page data field. - Reuse: no new auth path (passphrase, like config-retrieve + host-enroll); no new table (hub_settings); no new fetch credential downstream (the script reuses the config-retrieve git token).
v0.15.0 — Phase 2 managed updates: per-customer controller-version floor (2026-06-27)
The operator can now set a minimum controller version (FLOOR) — per-customer, defaulting to a global floor — and any box below it auto-updates to the floor on its next report (no customer action). The customer "update to latest" button is unchanged (latest, opt-in); the floor is the operator's enforced minimum and the auto-target (controlled rollout: the operator raises the floor).
store.go:customer_configsgainsmin_controller_version TEXT NOT NULL DEFAULT ''(per-customer override).- New
hub_settings(key,value)table holds the operator-set global floor (survives restarts). - Global default fallback
SetDefaultMinControllerVersion(config/envDEFAULT_MIN_CONTROLLER_VERSION). EffectiveMinControllerVersion(customerID)= per-customer override if non-empty, else global (hub_settings → config default), else "". PlusSet/GetGlobalMinControllerVersion,SetMinControllerVersion.
handler.go(handleReport): the controller report ACK (previously just{status, customer_blocked}) now also returns{min_controller_version: <effective floor>, latest_version: <registry latest>}. Both omitted when empty, so old controllers / unconfigured hubs behave exactly as before. NewSetLatestVersionProviderwires the registryVersionChecker(nil-safe).configs.go+ templates (operator UI, English): the Customers list shows a Floor column (effective floor, "override" tag, below-floor ● marker) + a global floor editor; each customer page shows the effective/global floor, below-floor status, and a per-customer override form. RoutesPOST /configs/global-floorandPOST /customers/{id}/floor(CSRF-protected; X.Y.Z or blank-to-clear).main.go: seeds the default floor from config/env; wires the latest-version provider.- Tests: effective-floor resolution (override beats global; empty when unset; DB-beats-default;
preserved across save); report ACK carries effective floor + latest + omits when unset; template render
smoke. Companion red-proof: breaking
EffectiveMinControllerVersionto return the global when an override is set makes the override-precedence store test AND the ACK test FAIL (verified, then restored). - Reuse: rides the existing report cycle (no new endpoint); the swap itself is the controller's Phase 1 flow + the agent — untouched here.
v0.14.0 — Passphrase-authed host enrollment (Day-0 option C) (2026-06-26)
Adds the single-secret Day-0 host-enrollment path proven in
documentation/audits/SPIKE-day0-firstboot-handshake-2026-06-26.md (option C). The operator /
host-bootstrap script now carries only the customer's retrieval passphrase — the global operator
key never enters the field deploy path.
- New endpoint
POST /api/v1/host-enroll(internal/api/handler.go,handleHostEnroll): passphrase-authed (X-Retrieval-Passwordheader, body{customer_id}), returns{host_id, api_key}. Mint-once-reuse — mints on first call (201), returns the existing credential byte-for-byte on every subsequent call (200), so re-running the bootstrap never orphans a running agent's key. Auth is checked before any mint (a wrong passphrase never writes a row): wrong/missing passphrase →401, unknown customer →404, missingcustomer_id→400. MirrorshandleConfigRetrieve's auth pattern +handleAdminCreateHost's mint block. - New store method
Store.GetHostByCustomer(internal/store/store.go):SELECT … FROM hosts WHERE customer_id = ? ORDER BY updated_at DESC LIMIT 1(usesidx_hosts_customer), nil-on-not-found. Backs the reuse lookup. >1 host for a customer (not expected in Day-0) → most-recent wins, never a duplicate mint. - Unchanged & deliberately untouched:
GET /api/v1/config/{id}(controller pull — same raw-YAML body) andPOST /api/v1/admin/hosts(global-key operator escape hatch, still PROVISIONAL pending the cutover lock-down). - Tests:
internal/api/host_enroll_test.go(mint/reuse/401-no-mint/404/400, +5) and aGetHostByCustomerstore test (+1); companion red-proof confirmed an always-mint variant fails the reuse assertion.
v0.13.1 — DR recipe v1 drive-shape sync: test-data + regression guard only (2026-06-16)
No behavior change — redeploy optional. Tracks the agent's v0.39.0 v1 host-half drive shape (which
dropped role + restic_repo_coord from drives[]). Because the hub reads drives as json.RawMessage
(verbatim passthrough), no store/handler/struct change was needed — only test-data + a regression guard.
internal/store/testdata/dr-recipe.golden.json+ thedrHostHalftest fixture: dropped therolekey fromdrives[0]to match the v1 shape the agent now emits.internal/api/testdata/host-report.golden.json: re-synced to be byte-identical with the agent'sinternal/hub/testdata/host-report.golden.json(sha25657f2a5e7…18b2f2b5). The hub copy previously lacked thedr_recipesection entirely; it is now a verbatim copy, so the cross-repo golden truly matches and POSTing it through/host-reportexercises theSaveDRRecipeHostHalfingest path.- New
TestAssembleDRRecipe_V1DriveShape— the regression guard: a stored host half whosedrives[]carry NEITHER dropped field but WHICH HAS apbsblock assembles cleanly (pbs carried through, drives passed through verbatim, neitherrolenorrestic_repo_coordpresent). Demonstrated to FAIL when the fixture re-addsrole, then reverted.
v0.13.0 — DR recipe: assemble + store + view the secret-free reconstruction recipe (2026-06-16)
DR recipe slice (hub half) — the assemble-store-view side of the secret-free reconstruction recipe
(documentation/audits/SPIKE-dr-recipe-2026-06-16.md). The hub receives two additive halves via the
existing report paths — the agent's storage/guest/PBS half (on the host-report) and the controller's
customer/apps half (on the controller report) — and assembles them into one operator-readable recipe per
customer. This is the clean inverse of the retired infra-backup: same "re-provision plan" goal, but
PLAINTEXT-at-rest is correct because the recipe has zero secrets.
- Store (
internal/store/dr_recipe.go): a DEDICATEDdr_recipetable (NOThost_escrow, NOT the droppedinfra_backup*tables) keyed bycustomer_id, holdinghost_half_json+app_half_json+recipe_version+host_id.SaveDRRecipeHostHalf/SaveDRRecipeAppHalfeach upsert their half and PRESERVE the other (last-write-wins per half).AssembleDRRecipestitches the two into anAssembledRecipe{recipe_version, customer, guests, pbs, drives, pve_storage, apps}— sub-sections pass through asjson.RawMessage(verbatim), ignore-unknown at the top level and version-skew tolerant (recipe_version= max of the two halves) for forward-compat across the three repos. - Ingest (
internal/api/handler.go):handleHostReportpersists thedr_recipehost-half (keyed by the host's customer);handleReportpersists thedr_recipeapp-half (keyed bycustomer_id) — mirroring theapp_telemetrypattern, backward-compatible (old agents/controllers omit the field), and never fatal to the heartbeat. - View (
internal/web/dr_recipe.go+ customer page): a DR-recipe panel on the customer detail page (which half has landed + last-updated) with a Download recipe (JSON) link →GET /customers/{id}/dr-recipe.jsonserves the assembled recipe (operator dashboard-auth, pretty JSON,Content-Dispositionattachment). No decrypt, nothing to redact. - Tests:
TestDRRecipe_StoreRoundTrip(each half preserves the other),TestAssembleDRRecipe_MatchesGolden(the assembled wire shape pinned intestdata/dr-recipe.golden.json),TestAssembleDRRecipe_IgnoreUnknownAndVersionSkew(a forward-compat half still assembles; version = max),TestAssembleDRRecipe_PartialHalves(one half present),TestAssembleDRRecipe_NoSecrets(defense-in-depth credential-key sweep). Pairs with felhom-agent v0.38.0 + felhom-controller v0.73.0.
v0.12.0 — retire Infra Backup + purge its plaintext secrets + fix the daily backup-deadline email (2026-06-16)
Phase-1 of the Infra Backup retirement (per documentation/audits/SPIKE-infra-backup-2026-06-15.md).
The mechanism had been dead since slice 8C, yet the hub still stored each version as a plaintext
JSON blob at rest containing the customer's app-secret encryption key, restic password, and
Cloudflare tokens — a zero-knowledge violation. Its absence was also the root cause of the daily
expected_backup_missed false-alarm email.
Changed
- Backup-deadline check repointed to PBS freshness.
monitor.CheckBackupDeadlinesno longer looks for abackup_completedevent (no component emits it anymore — the disk-tier backup moved to the agent in slice 8C, so the check fired daily for every healthy customer). It now reads the customer's latest agent host-report and raisesexpected_backup_missedonly on positive evidence: no PBS snapshot / successful vzdump at all, the newest backup older than 26h, or the newest PBS snapshot'sverify_state == "failed". A fresh-but-not-yet-verified snapshot is not a failure (PBS verifies on its own cadence) — alarming on it would just re-create the false alarm. The db-dump half is unchanged (the in-guest controller still emitsdb_dump_completed). A customer with no host-report (legacy/defunct) gets no backup alarm here — liveness is the host-staleness checker's job. New store accessorGetLatestHostReportJSON. Tests:internal/monitor/deadline_test.go(fresh+verified→quiet, stale→alarm, failed-verify→alarm, no-report→quiet, db-dump half preserved, plus a pureassessBackupFreshnesstable). The fresh+verified→quiet test is the companion: it fails against the old event-based check.
Removed
- The Infra Backup feature: ingest endpoint
POST /api/v1/infra-backup, gettersGET /api/v1/infra-backup/{id}[/versions]and their handlers; store methodsSaveInfraBackup/GetInfraBackup/GetInfraBackupByID/GetInfraBackupMeta/ListInfraBackupVersions/pruneInfraBackups+ theInfraBackupMeta/InfraBackupVersiontypes; the operator "Infra Backup" panel (customer_unified.html,customer.html). TheGET /api/v1/recovery/{id}endpoint is kept but now returns only the generatedconfig_yaml(no infra-backup payload). The customer-page config-drift badge that diffed against the stored controller.yaml is hidden (its at-rest source is gone); the live "Show Diff" path is unaffected.
Security / migration
- Plaintext secret purge.
migrate()nowDROPsinfra_backup_versions+infra_backupsand runsVACUUM(+wal_checkpoint(TRUNCATE)) so the freed pages holding the plaintext keys/ tokens are physically reclaimed, not merely delinked. Gated on table existence so normal restarts don't pay the VACUUM cost. - Out of scope (flagged for the operator): the exposed Cloudflare / hub / session credentials in
the dropped blobs remain valid until rotated (operator step). Separately, the legacy
reportstable holds thousands of historical rows with a plaintextrestic_passwordvalue from old controller versions — a distinct leak, not purged here (the live controller no longer sends it).
v0.11.0 — slice 10D: DR capstone — recovery mode + re-enroll + directive serving (2026-06-10)
The hub half of the slice-10 DR capstone (closes slice 10). The hub ORCHESTRATES recovery but holds
no usable secret and no Cloudflare write-power: the escrow blobs it serves are opaque (need R,
which the hub never has), and the destructive tunnel/PBS rotation is the operator's step from a
trusted environment. A compromised hub can at most hand out opaque blobs + rotate/revoke its own
per-host credential — it cannot hijack a customer's tunnel.
Added
PUT /admin/hosts/{id}/recovery-mode(global key) — arm recovery mode with a bounded TTL (ttl_seconds, clamped [60s, 4h], default 30m → auto-expires);DELETEto disable. The restore directive + re-enroll are served ONLY while recovery mode is active.POST /hosts/{id}/re-enroll— gated ONLY on recovery mode (the lost box has no old key; the operator armed recovery mode after out-of-band validation). Rotates the host's API key to the new box's key (the old box's hub access is revoked instantly) and returns the DR directive + the two opaque escrow blobs. Without recovery mode → 403. Zero-knowledge: even a wrongful re-enroll in the window leaks nothing recoverable (the blobs needR).GET /hosts/{id}/restore-directive(re-enrolled key, recovery-gated) — re-fetch the directive.- Store/escrow:
hosts.recovery_mode_until(additive);host_escrow.identity_blob+directive_json(the age-wrapped identity blob + non-secret directive, stored alongside the K-escrow). Methods:SetRecoveryMode/ClearRecoveryMode,RotateHostAPIKey,SaveHostDRBundle/GetHostDRBundle. The slice-7 escrow upload (PUT /hosts/{id}/escrow) now also acceptsidentity_blob_b64+directive(additive).
Not built (by design — the locked rotation model)
- No Cloudflare write-credential in the hub. The operator deletes the stale tunnel connector + rotates the tunnel/PBS token from their trusted environment (a documented procedure / future small operator CLI). The hub may optionally hold a read-only CF token to surface connector state.
Tests
- re-enroll refused without recovery mode (403); recovery-mode arm is global-key-only; re-enroll rotates + revokes (old key → 401, new key → 200); directive served only in recovery mode + expires; clear disables re-enroll.
v0.10.0 — slice 10B: signed-op job completion (clear-job) (2026-06-10)
The hub half of slice 10B is small by design — the hub stores + serves the operator-signed blobs opaquely (it holds no signing key and can neither forge nor open them; the agent verifies + executes). 10B adds the missing completion path so a processed job leaves the queue.
Added
DELETE /api/v1/hosts/{host_id}/jobs/{job_id}(per-host key, self-scoped; the global key may clear any) — the agent calls it after executing OR terminally rejecting a job. Idempotent (clearing an absent job is a clean 200). Store:DeleteSignedJob.
Unchanged (already in 10A, reused by 10B)
POST /admin/hosts/{id}/jobs(operator enqueues the signed blob),GET /hosts/{id}/jobs(the agent fetches), and thehas_signed_opsenvelope flag. The signed blob stays opaque on the wire (a base64{op_blob_b64, sig_armored}envelope the agent parses) — no jobs-wire golden change.
Tests
DELETE …/jobs/{id}is self-scoped (host A cannot clear host B's job → 403) and idempotent.
v0.9.0 — slice 10A: desired-state serving + signed-jobs queue (the "Down" channel) (2026-06-10)
The hub half of slice 10A: the hub now serves operator intent down to already-authenticated
hosts. The control envelope (the host-report response) stops returning placeholder
desired_generation:0 / has_signed_ops:false and carries the host's real generation + a
signed-jobs flag — the cheap change-notification the agent (v0.15.0) acts on. The heavy
desired-state moves only on a dedicated, self-scoped fetch.
Added
PUT /api/v1/admin/hosts/{host_id}/desired-state(global/operator key only) — sets a host's desired-state and atomically bumpsdesired_generation. The body is JSON the hub stores + serves opaquely (it validates only that it is well-formed JSON; the agent/CLI owns the schema). Unknown host → 404; malformed JSON → 400. Minimal admin path; rich editing UX is later.GET /api/v1/hosts/{host_id}/desired-state(per-host key, self-scoped — a host reads only its own; the global key may read any) — returns{generation, desired_state}. The agent fetches it when the envelope's generation advances past its cache.GET /api/v1/hosts/{host_id}/jobs(per-host key, self-scoped) — serves the host's pending opaque signed-op blobs (oldest first). The hub never forged, opened, or executes them (verify + run is slice 10B; this only serves the queue).POST /api/v1/admin/hosts/{host_id}/jobs(global key only) — enqueues a pre-signed opaque job blob. The minimal operator path to seed the queue; the hub holds no signing key.- Store: a new
signed_jobstable (per-host opaque blob queue);SetHostDesired(set + bump generation, atomic),EnqueueSignedJob/GetSignedJobs/CountSignedJobs. Thehoststable's previously-inertdesired_json/desired_generationcolumns are now live.
Changed
- The host-report control envelope now reports the host's actual
desired_generationandhas_signed_ops(queue non-empty), both degrading safely to their old defaults on a store error (a heartbeat never fails on the control channel).poll_interval_seconds/blockedunchanged.
Tests
- admin-set bumps the generation each write + the served state reflects the latest body; admin-set is global-key-only (per-host → 403, malformed → 400, unknown host → 404).
GET /desired-stateis self-scoped (host A's key → host B → 403; global → any; no token → 401).- the envelope carries the current generation +
has_signed_opsflips on enqueue;GET /jobsis self-scoped + serves the blobs oldest-first; admin enqueue is global-key-only. - cross-repo golden round-trip:
testdata/desired-state.golden.jsonset → fetched back unchanged (the opaque pass-through), byte-identical with felhom-agent's copy.
(no version bump) — slice 9 cross-repo wire-contract: host.cpu_temp_c (2026-06-10)
Slice 9 adds a nullable cpu_temp_c field to the shared HostMetrics wire struct (the agent's
new CPU/chassis-temperature collector). The agent's host-report carries it too, so the hub's
cross-repo host-report golden (internal/api/testdata/host-report.golden.json) was updated to
stay byte-identical with felhom-agent/internal/hub/testdata/host-report.golden.json (the
duplicated-contract discipline; manual diff confirmed identical). No hub code change — the full
report_json already persists the field verbatim, and the hub does not surface CPU temp on the
operator dashboard yet (an optional later freebie). The golden-contract test (host_test.go) still
passes (the host parse-struct ignores the extra key).
v0.8.0 — opaque PBS recovery-code escrow storage (slice 7, doc 03 §8a) (2026-06-10)
Hub half of slice-7 close-out: store the agent's opaque R-wrapped PBS-key escrow blob. The
default posture is zero-knowledge — the hub holds ciphertext it cannot open (it has no recovery
code; there is no decrypt path). Pairs with felhom-agent v0.9.0 (escrow creation). Consumption /
restore-mode serving is slice 10.
Added
PUT /api/v1/hosts/{host_id}/escrow— authed with the per-host key (a host may only write its own escrow; the global operator key is also accepted). Body mirrors the agent's emit struct (blob_b64,key_fingerprint,posture,created_at). Stores the decoded opaque bytes verbatim; rotation is last-write-wins. No serving this slice.host_escrowtable (host_idPK,blobBLOB, fingerprint/posture/created_at). Store methodsSaveHostEscrow/GetHostEscrow(HostEscrow). The hub never transforms or decrypts the blob.
Tests
- Stores the opaque blob verbatim (round-trips byte-identical); rotation last-write-wins; rejects an absent/wrong key (401) and a host writing another host's escrow (403); bad/empty base64 → 400; the wire-contract key-set matches the agent's emit struct.
Security note
The hub stores ciphertext only — holding the blob does NOT let Felhom read customer data (separation principle, doc 03 §8a). The per-host-key gate scopes writes to the owning host.
v0.7.5 — restore-test "passed with warnings" visibility (2026-06-09)
Hub half of TASK — Restore-test must not false-fail on benign start warnings (Phase B). The
agent (v0.7.0) now treats a guest-start advisory like the systemd-nesting warning as a PASS
(verdict is liveness, not the start exitstatus) and carries the warning text on the wire. This
makes that visible to the operator instead of indistinguishable from a clean pass.
Added
hostRestoreTest.warnings([]string) +warnings_recognized(bool) mirror fields, matching the agent'shub.RestoreTestwire contract (omitempty; an absentwarnings_recognized⇒false⇒ treated as the louder unrecognized case — a missing flag can only over-notice).
Changed
- Host-report ingest now surfaces a passed restore-test that carried warnings:
[INFO] restore-test passed WITH WARNINGS (recognized)when every warning is the known-benign anchor, escalated to[WARN] … UNRECOGNIZED WARNINGSotherwise — as loud as a failed PBS verify, so a real restore warning can't hide behind a green pass. A FAILED restore-test still logs the existing[WARN] … FAILED.
Tests / contract
restore_tests[0]in the host-report golden gainswarnings+warnings_recognized; the golden stays byte-identical with felhom-agent's copy (sha256-verified) and the bidirectional key-set contract test now round-trips the new keys throughhostRestoreTest.
Not in this slice
- No dashboard widget: the hub web layer renders only controller-report data — there is no host-domain dashboard surface yet (guests/storage/restore_tests/pbs_snapshots are log+persist only, same as the failed-PBS-verify signal). Distinct dashboard treatment lands when the host-domain dashboard does (slice 10). The operator signal this slice is the log line.
v0.7.4 — ingest agent pbs_snapshots (slice 6 Phase B) (2026-06-09)
The agent's slice-6 Phase B work populates the host-report's pbs_snapshots (the PBS offsite
inventory + per-snapshot verify-state). This is the hub half: accept + persist them. Minimal —
the rich offsite policy is hub-owned (slice 10); this mirrors what the agent reports.
Added
hostPBSSnapshotmirror struct inhostReportPayload(internal/api/handler.go) — field-for-field with the agent'shub.PBSSnapshotwire contract (namespace/backup_type/ backup_id/backup_time/size_bytes/owner/protected/encrypted/verify_state/verify_upid). Persisted viareport_json(no new columns — the slice-5/6A precedent).- A FAILED PBS verify is logged prominently (
[WARN]— the loudest offsite-DR signal, same treatment as a failed restore-test). Thehost-reportinfo line now counts pbs-snapshots. testdata/host-report.golden.jsonupdated with a populatedpbs_snapshots[0], kept byte-identical with felhom-agent's copy.TestHostPBSSnapshot_GoldenContract— the hub half of the bidirectional key-set test.
Notes
- Backward-compatible: an agent that omits/empties
pbs_snapshotsis accepted unchanged.
v0.7.3 — ingest agent backups + restore_tests (slice 6 Phase A) (2026-06-09)
The agent's slice-6 work populates the host-report's backups + restore_tests (the
self-restore-test result). This is the hub half: accept + persist them. Minimal — the rich
backup policy (schedule/retention/target selection) is hub-manifest-owned and lands at
slice 10; this slice only mirrors what the agent reports.
Added
hostBackup/hostRestoreTestmirror structs inhostReportPayload(internal/api/handler.go) — field-for-field with the agent'shub.Backup/hub.RestoreTestwire contract. Persisted verbatim inreport_json(no new columns — slice-5 precedent).- A FAILED restore-test is logged prominently (
[WARN], the loudest DR signal there is); a failed backup is logged too. Thehost-reportinfo line now counts backups + restore-tests. testdata/host-report.golden.jsonupdated with a populatedbackups[0]/restore_tests[0], kept byte-identical with felhom-agent's copy.TestHostBackup_GoldenContract/TestHostRestoreTest_GoldenContract— the hub half of the bidirectional key-set test (round-trip the golden through the mirror, assert exact keys).
Notes
- Backward-compatible: an agent that omits/empties these is accepted unchanged. The legacy controller report path is untouched (frozen until slice 10).
v0.7.2 — ingest agent storage_targets (slice 5 Phase A) (2026-06-09)
The agent's slice-5 work populates the host-report's storage_targets (previously empty).
This is the hub half: accept + persist them. Minimal by design — the rich, authoritative
storage manifest (desired class/role/policy/creds) is hub-owned and lands at slice 10; this
slice only mirrors what the agent observes.
Added
hostReportPayload.StorageTargets(internal/api/handler.go) — a full mirror of the agent'shub.StorageTargetwire contract (name/type/durable_id/state/reachable/usage/ content/mount/class_hint/role/thin_pool/smart). The targets are persisted verbatim in the existingreport_jsonrow (no schema change); the handler counts them and logs a[WARN]when any aredisconnected(the storage analog of host-down visibility).testdata/host-report.golden.json— updated to carry two populatedstorage_targets(an lvmthin withthin_pool, a usb), kept byte-identical with felhom-agent's copy.TestHostStorageTarget_GoldenContract— the hub half of the bidirectional key-set test: round-trips the golden'sstorage_targets[0]through the mirror struct and asserts the key set matches exactly (no missing/extra fields vs the agent).TestHostReport_GoldenContractalso now asserts the targets are persisted + parse back.
Notes
- Backward-compatible: an older agent that sends
storage_targets: [](or omits it) is accepted unchanged. The legacy controller report path is untouched (frozen until slice 10).
Repo docs — no hub version change (2026-06-08)
Changed
- Reflowed
felhom.eu/CLAUDE.md— removed hard mid-paragraph line wraps (prose, list items, blockquotes now single-line); tables untouched; rendered output unchanged. - Unified the REPORT/CHANGELOG convention: this repo's
REPORT.mdswitches from append/cumulative to overwrite-latest (uniform with the sibling repos);CHANGELOG.md(this file) stays the cumulative log, newest on top. UpdatedREPORT.md's header note accordingly (existing sections retained as history). Added an explicit no-secrets rule. No hub code change → no version bump.
v0.7.1 (2026-06-08)
Changed
/host-reportrejects oversize bodies explicitly with 413 (handler.go) instead of silently truncating at the 4 MiBLimitReadercap. Reads one byte pastmaxHostReportBytesand returns413 Payload too large— a truncated-but-valid JSON could otherwise be accepted as a partial report (silently dropping guests from the mirror). The controllerhandleReport1 MiB path is unchanged (frozen until slice-10 cutover).
Added
- Cross-repo contract fixture
hub/internal/api/testdata/host-report.golden.json(byte-identical with felhom-agent's copy) +TestHostReport_GoldenContract— POSTs the golden through the realhandleHostReportand asserts 200 + denorm (guest_total/guest_running/cloudflared_status) + both guests upserted, provinghostReportPayloadstill extracts the contract from the real shape. Duplicated contract (no shared types module yet); revisit at slices 5/6.
v0.7.0 (2026-06-08)
Added — host-domain ingest (slice 3, additive; controller path untouched)
- New tables
hosts,guests,host_reports(store.go migrate(), idempotent). Full schema now, including columns inert until slice 10 (hosts.desired_json/desired_generation/dr_record_json,guests.api_key/desired_spec_json) so the cutover needs noALTER. Nothing reads/writes the inert columns this slice. POST /api/v1/host-report— the agent's heartbeat. Per-host Bearer auth; 4 MiB body; persists the full report + denormalized fields (cpu/mem/disk %, guest counts, cloudflared status); upserts each guest's reality columns (guest_id = "<host_id>/<vmid>", hub-derived); returns the control envelope{status, poll_interval_seconds:900, blocked, desired_generation:0, has_signed_ops:false}(blockedreflects the customer's status; the latter two are reserved/placeholder for slice 4).- Per-host key auth —
checkAuthHost(Bearer → host → customer), added alongside the unchangedcheckAuthCustomer. Global key remains a bootstrap fallback. POST /api/v1/admin/hosts— PROVISIONAL global-key-only host mint (host_id + per-host api_key); the slice-3 bootstrap until enrollment (slices 7–8) replaces it.- Host dead-man's-switch —
monitor.HostStalenessCheckeroverhost_reports, emittinghost_stale/host_down/host_recovered(30m/60m), attributed to the host's customer; registered inallowedEventTypes; wired incmd/hub/main.goon the existing 60s ticker. A deliberate sibling of the controllerStalenessChecker(both run until slice 10). - Store methods:
GetHostByAPIKey,GetHost,ListHosts,UpsertHost,SaveHostReport,UpsertGuestFromReport(preserves inert columns on conflict),GetHostStaleness(skips never-reported hosts),GuestID.Prunenow also pruneshost_reports(same retention). - Tests (new, hermetic): store, auth (
checkAuthHost), ingest (valid+envelope+denorm, host_id mismatch→403, unknown-host-under-global→400, blocked→true, oversize→400), admin mint (non-global→403, unknown customer→400, mint+round-trip), host staleness transitions.
Unchanged (explicit)
- The controller path —
/api/v1/report,reports,customer_configs,checkAuthCustomer, the existing staleness/deadline checkers — is untouched and still green. The old controller and the new agent report in parallel during slices 3–9; the schema/auth cutover is slice 10.
v0.6.2 (2026-02-26)
Added
- Infra backup GFS retention — New
infra_backup_versionstable stores multiple backups per customer. GFS pruning keeps: all from last 24h, latest per day (7 days), latest per week (4 weeks), latest per month (3 months) — ~14 versions max per customer GET /api/v1/infra-backup/{id}/versions— Returns metadata list of all retained backup versions (date, stack names, disk count) for a customer. Bearer auth.- Recovery version selection —
GET /api/v1/recovery/{id}?version=IDfetches a specific backup version instead of latest. Response now includesbackup_versionsarray with all available versions. - Dashboard backup history — Customer detail page "Infra Backup" card shows version count and collapsible history table (date, apps, disks)
Changed
SaveInfraBackup()— Now INSERTs a new row instead of upserting, preserving history. Automatically prunes old versions via GFS algorithm.- One-time migration — Existing data from
infra_backupstable is copied toinfra_backup_versionson first startup
v0.6.1 (2026-02-25)
Added
- Delete issues from app detail page — Known Issues table now has per-row checkboxes with "Delete Selected" and "Delete All Issues" buttons; keeps telemetry data (memory trends, etc.) intact
DELETE /apps/{appName}/delete-issues— New POST endpoint supportingaction=selected(withissue_idsform values) andaction=all
Fixed
- Hub-side fingerprint hardening —
fingerprintIssue()now strips ANSI escape codes, ISO/syslog timestamps, and lowercases before truncating to 100 chars. Prevents duplicate issue rows when messages differ only by embedded timestamps.
v0.6.0 (2026-02-25)
Added
- Geo-restriction display (
customer_unified.html) — New "Geo-korlátozás" section on customer detail pages showing: enabled/disabled status, allowed countries, per-app overrides, last sync time, and sync errors. Only visible when the controller reports geo_restriction data. - "Összes geo-korlátozás eltávolítása" button — One-click removal of all
[felhom-geo]Cloudflare WAF rules. The Hub calls the Cloudflare API directly (bypasses potentially blocked tunnel), then retries notifying the controller in background (every 30s for up to 10 min) to disable geo in its settings. - Cloudflare unblock client (
internal/cloudflare/unblock.go) — Minimal Cloudflare API client for deleting geo-restriction WAF rules. Resolves zone ID, finds thehttp_request_firewall_customruleset, and deletes rules with[felhom-geo]description prefix. POST /customers/{id}/geo/disableroute — CSRF-protected endpoint for the geo-disable action.
Removed
- Legacy Monitoring UUIDs — Removed the "Monitoring UUIDs" section from the config form (
config_form.html), UUID form-field handling frombuildConfigJSON(), UUID import fromhandlePullConfig(), volatile key entries formonitoring.ping_uuids.*, and the commented-outping_uuidssection fromcontroller.yaml.default. Monitoring is fully handled by the Hub event system since v0.3.0.
v0.5.0 (2026-02-25)
Added
- Configuration page (
GET /configuration) — New "Configuration" tab in the web UI with asset management controls. Displays asset file count, manifest generation timestamp, and a "Refresh Assets from Image" button. - Manual asset re-seed (
POST /configuration, action=refresh_assets) — Re-reads the baked-in seed directory, compares SHA-256 checksums with PVC assets, and updates changed files. Rebuilds the manifest afterward. Controllers pick up changes on their next daily sync. ReSeed()method (internal/assets/assets.go) — Public method for triggering asset re-seed + manifest rebuild from the web UI.
Changed
- Asset seeding:
seedIfEmpty()→seedOrUpdate()(internal/assets/assets.go) — On startup the Hub now compares SHA-256 checksums between the image seed directory and the PVC, updating any changed files instead of only seeding into an empty directory. This means redeploying the Hub image with updated assets automatically propagates them without PVC deletion. isAssetFile()expanded — Now also matches*-favicon.svgand*-favicon.icopatterns, allowing branding assets likefelhom-favicon.svgin the manifest.RebuildManifest()refactored — Internal logic extracted torebuildManifestLocked()for reuse byReSeed().- Web Server struct — Added
assetsMgrfield andSetAssetManager()method. Wired inmain.go. - All templates translated to English — The "Alkalmazások" nav link and telemetry pages (apps.html, app_detail.html, customer_unified.html telemetry section) are now in English, consistent with the rest of the Hub UI.
- Navigation updated — All templates now show four tabs: Dashboard, Customers, Apps, Configuration.
v0.4.1 (2026-02-23)
Added
- Per-app telemetry reset (
store/telemetry.go,web/apps.go) — New "Telemetria törlése" button on the app detail page that deletes all telemetry records and known issues for the selected app. Useful after major app updates when old data is no longer representative. Includes confirmation dialog and flash notification. DeleteAppTelemetry()andDeleteAppIssues()store methods (store/telemetry.go) — Delete all telemetry/issue rows for a specific app_name.POST /apps/{name}/reset-telemetryroute (web/server.go) — CSRF-protected endpoint that triggers the reset and redirects back with flash message.
v0.4.0 (2026-02-23)
App Telemetry & Analytics Dashboard
Added
app_telemetryandapp_log_issuesSQLite tables (store/store.go) — store per-app resource metrics and deduplicated log issues reported by v0.28.0+ controllers.internal/store/telemetry.go— New store methods:SaveAppTelemetry,GetFleetAppSummary(with P95 memory calculation),GetAppTelemetryHistory,GetAppCustomerBreakdown,GetCustomerAppSummary,GetAppIssues,GetRecentIssuesAllApps,PruneAppTelemetry,PruneStaleIssues. New types:AppTelemetryRecord,FleetAppSummary,AppTelemetryPoint,AppCustomerStats,CustomerAppSummary,AppIssue./api/v1/reporthandler update (api/handler.go) — After saving the standard report, parses the optionalapp_telemetryJSON field and persists it. Backward-compatible: old controllers (noapp_telemetrykey) are unaffected.- Fleet app list page (
GET /apps) — Hungarian-language dashboard showing all deployed apps fleet-wide with deployment count, avg/P95 memory, catalog estimate/limit accuracy, error/warning badges. Sortable columns, 24h/7d/30d period selector. - Per-app detail page (
GET /apps/{name}) — Memory trend Chart.js chart (avg + peak, with catalog limit line), per-customer breakdown table, known log issues table (severity, message, occurrence count, affected customers). Includes suggested mem_limit from P95×1.2 rounded to 32M. - Customer detail page telemetry section (
customer_unified.html) — New "Alkalmazás telemetria" card with per-app memory (current/avg/peak) and log error/warning counts linking to /apps/{name}. - Chart.js (
static/chart.min.js) — Embedded from controller build, served at/static/chart.min.js. - "Alkalmazások" nav link — Added to header navigation across all templates.
- New CSS (
style.css) —.badge,.badge-error,.badge-warn,.summary-cards,.summary-card,.chart-container,.period-selector,.period-btn,.accuracy-dot,.mem-ok/warn/danger,.data-tablestyles. - Telemetry pruning (
cmd/hub/main.go) —pruneAll()now also prunes app_telemetry rows older than 90 days and stale log issues not seen in 30 days.
Changed
internal/web/apps.go(new file) —handleApps,handleAppDetail,parsePeriod,sortFleetSummary,aggregateHistoryForChart,parseLimitMB,memoryColor,accuracyClass,getCSRFTokenhelper functions.internal/web/server.go— Added routes for/apps,/apps/{name},/static/chart.min.js. AddedmemoryColor,accuracyClass,gttemplate functions.internal/web/embed.go— Added//go:embed static/chart.min.jsdirective.
v0.3.7 (2026-02-21)
Asset management API
- New
internal/assetspackage: manages app assets (logos, screenshots) on Hub PVC (/data/assets/) with automatic seeding from baked-in image copy on first run. - Two new authenticated API endpoints for controllers to sync assets:
GET /api/v1/assets/manifest— returns JSON manifest with filenames + SHA-256 checksumsGET /api/v1/assets/file/{filename}— serves individual asset files
- Dockerfile updated to
COPY assets/ /usr/share/felhom/assets-seed/for first-run seeding. - Build script syncs website assets (
*-logo.{svg,png},*-screenshot-*.webp) into Docker build context.
v0.3.6 (2026-02-21)
Human-friendly retrieval passwords
- Retrieval passwords now use Hungarian word passphrases (e.g.
áldás-plazmid-palánta-süvítve-pócgém) instead of 64-char hex strings. - Embedded 29K+ curated Hungarian word list (
hungarian.txt) via go:embed; 5-word passphrases give ~74 bits of entropy. - New
configgen.RandomPassphrase(wordCount)function; all 3 retrieval password generation sites updated. - API keys remain as hex (machine-to-machine, never typed by humans).
v0.3.5 (2026-02-21)
Recovery Endpoint & Customer Standing
- New
GET /api/v1/recovery/{customer_id}endpoint: returns both generated controller.yaml and infra backup in a single response for disaster recovery. Auth viaX-Retrieval-Passwordheader (same as config retrieval). - Report response now includes
customer_blocked: truewhen customer status is "blocked" — allows controllers to detect standing and enter limited mode.
v0.3.4 (2026-02-20)
- Rename version labels: "Current version" → "Controller version", "Latest version" → "Registry latest".
v0.3.3 (2026-02-20)
Bugfixes
- Fix double "v" prefix in controller version display (showed "vv0.21.1" instead of "v0.21.1").
- Skip deprecated
monitoring.ping_uuids.*keys in config diff comparison (added to volatile keys).
v0.3.2 (2026-02-20)
Hub Version Display
- Show Hub version in footer of all pages via
hubVersiontemplate function. web.New()now acceptsversionparameter (4th arg) — set via ldflags at build time.
v0.3.1 (2026-02-20)
Config Diff Display + Pull Config
- Value-based config comparison: Replaced broken SHA256 hash comparison with semantic YAML comparison. Both configs are parsed into maps, flattened to dot-notation keys, and compared by value. Ignores key ordering, whitespace, comments, and volatile fields (
web.session_secret). Shows actual diff count on customer page ("⚠ Config mismatch — N differences"). - Config diff endpoint (
GET /customers/{id}/config-diff): Fetches live YAML from controller via newGET /api/configendpoint, generates Hub YAML viaconfiggen.Generate(), returns JSON with per-key diffs (key, hub value, controller value, status). Sensitive values (tokens, passwords, secrets) are masked. - Pull Config (
POST /customers/{id}/pull-config): Reverse of Push Config — imports controller's current config into the Hub. Extracts identity fields (name, domain, email) and override fields (infrastructure tokens, git credentials, monitoring UUIDs). Preserves existing APIKey and RetrievalPassword. - Diff display UI: "Show Diff" button on customer page expands a table showing all key-value differences with color-coded rows (yellow=changed, blue=hub-only, orange=controller-only).
- Pull Config button: Added next to existing "Push Config" with confirmation dialog.
v0.3.0 (2026-02-20)
Hub Monitoring Takeover — Event System, Dead Man's Switch, Notifications
Replaces external Healthchecks.io with a Hub-native event system. The Hub becomes the single source of truth for all customer monitoring, event tracking, dead man's switch alerting, and notification delivery.
Phase 1 — Event System
eventstable in SQLite: stores all events with customer_id, event_type, severity, message, details_json, source, timestamp- Indexes:
idx_events_customer_created(customer + time DESC),idx_events_type(type + time DESC) - Store methods:
SaveEvent,GetRecentEvents,GetEventsByType,GetLatestEventByType,GetAllRecentEvents,CountEventsBySeverity,PruneEvents,GetActiveCustomerIDs POST /api/v1/eventendpoint: accepts structured events from controllers, validates event_type against 27 allowed types, validates severity (info/warning/error), stores in DB- Enhanced auth:
checkAuthCustomer()validates per-customer API keys match the customer_id in payload; global key bypasses ownership check - Prune: events pruned alongside reports at 04:30 Budapest time
Phase 2 — Dead Man's Switch
- Staleness checker (
internal/monitor/staleness.go): runs every 60s, detects when controllers stop reporting- ok→stale (>30min): inserts
node_stalewarning event - any→down (>60min): inserts
node_downerror event - stale/down→ok: inserts
node_recoveredinfo event - Skips blocked customers, no false alerts on startup
- ok→stale (>30min): inserts
- Backup deadline checker (
internal/monitor/deadline.go): runs daily at 05:00 Budapest- Detects missing
backup_completedevents since midnight → insertsexpected_backup_missederror - Detects missing
db_dump_completedevents → insertsexpected_dbdump_missederror - Grace: skips customers with
node_downstate
- Detects missing
scheduleDaily()helper: goroutine that sleeps until target time (Europe/Budapest), runs function, loops/healthzenhanced: returns 503 if SQLite Ping fails
Phase 3 — Notification System
- Dispatcher (
internal/notify/dispatcher.go): processes events and sends emails via Resend API- Operator channel: English emails to operator for warning/error events, 1h cooldown per customer:eventType
- Customer channel: Hungarian emails per event_type, respects customer preferences (enabled_events, cooldown_hours), blocked customers skipped
- Test bypass:
testevent type skips cooldown/preferences, sends directly to customer email
- Email templates (
internal/notify/templates.go): operator (concise English), customer (Hungarian per event type with complete message table) - Cooldown tracking: in-memory maps with per-customer:eventType granularity
customer_notificationstable: addedcooldown_hourscolumn (default 6)notification_logtable: addedchannelcolumn (operator/customer)- Wired into
/api/v1/eventhandler and staleness/deadline checkers
Phase 4 — Hub UI
- Events section on customer detail page: last 50 events, severity filter buttons (All/Errors/Warnings/Info), colored severity badges
- Dashboard badges: error+warning count in last 24h per customer, clickable to customer events
- Notification log: shows channel column (operator/customer) in customer detail page
- Config form: Monitoring UUIDs section marked as "Legacy" with deprecation notice, collapsed by default
Phase 6 — Config Cleanup
controller.yaml.default:monitoring.ping_uuidssection commented out (deprecated)buildConfigJSON: only writesping_uuidsto config JSON if user explicitly provides UUID values (new configs get none)
v0.2.2 (2026-02-20)
Config Hash Comparison
- Config sync status on unified customer page: compares SHA256 hash of controller's
controller.yaml(from report payload) against Hub-generated YAML. Shows "In sync", "Config mismatch", or "Unknown" (controller needs v0.20.0+ to report hash). - Visible in the Controller Update section next to Push Config button.
v0.2.1 (2026-02-20)
Unified Customer Management
All customer views consolidated into a single page. New management features: blocked status, dashboard merge, config push, and auto-config creation.
New features
-
Unified customer page —
/customers/{id}:- Single page showing both configuration info and live report data
- Replaces separate
/configs/{id}(config detail) and/customers/{id}(report detail) pages - Shows config management (credentials, setup commands, YAML preview) when config exists
- Shows "Create Config" button for manual (report-only) customers
- Old
/configs/{id}URLs redirect to/customers/{id}
-
Dashboard shows pending customers:
- Customers with config but no reports appear on dashboard with "PENDING" status
- All metric columns show "—" for pending customers
-
Blocked/Banned status:
- Customers can be blocked via button on detail page
- Blocked customers hidden from Dashboard
- Reports still accepted (prevents controller retry loops) but notifications suppressed
- "BLOCKED" badge shown on Customers list and detail page
- One-click unblock button
-
Config push to controller:
- "Push Config" button on unified page (visible when controller URL known)
- Generates YAML and POSTs to
{controller_url}/api/config/apply - Note: requires controller v0.20.0+ with config apply endpoint
-
Auto-create config from report data:
- "Create Config" button on manual customer pages
- Pre-fills customer name from report, generates credentials
- Redirects to edit form for additional fields
Changes
- Customers list: all rows now link to
/customers/{id}(unified page) - Config badges: new MANAGED/MANUAL/BLOCKED pill-style badges
customer_configstable: addedstatuscolumn (active/blocked)- Status functions handle "pending" and "blocked" status values
v0.2.0 (2026-02-20)
Customer Configuration Management
New "Configurations" section for pre-provisioning customer nodes. Operators can configure
customer settings in the Hub web UI, then docker-setup.sh downloads a ready-made
controller.yaml — reducing deployment to a customer ID and password.
New features
-
Web UI —
/configspages:- List all customer configurations in a table
- Create new configuration: customer identity, infrastructure secrets (CF tunnel/API tokens), git sync credentials, monitoring UUIDs — organized in collapsible sections
- Detail page: shows credentials (retrieval password, per-customer API key) with copy-to-clipboard,
setup commands (
docker-setup.shandcurl), live YAML preview - Edit and delete configurations
- Navigation tabs (Dashboard / Configurations) on all pages
-
Config retrieval API —
GET /api/v1/config/{customer_id}:- Authenticated via
X-Retrieval-Passwordheader (separate from Bearer token) - Generates complete
controller.yamlby deep-merging template with customer overrides - Template sourced from
controller.yaml.example(fetched from Gitea repo periodically) - Falls back to embedded default template if fetcher not configured
- Authenticated via
-
Per-customer API keys:
- Each customer config gets its own API key (auto-generated, 64 hex chars)
- Controllers can authenticate with per-customer key instead of the shared global key
- Backward compatible — global
report_api_keycontinues to work alongside per-customer keys
-
YAML generation (
internal/configgenpackage):- Deep-merge of template + customer-specific overrides
- Programmatic injection: customer identity, hub config, session secret
- Shared by both API handler and web UI preview
-
Template fetcher (background goroutine):
- Periodically fetches
controller.yaml.examplefrom Gitea (configurable interval) - Requires
registry.username+registry.tokenin hub.yaml - Falls back to
go:embeddefault template when not configured
- Periodically fetches
-
Data layer:
- New
customer_configsSQLite table - 6 CRUD methods: Save, Get, List, Delete, GetByAPIKey, UpdateRetrievalPassword
- New
Configuration
New registry section in hub.yaml:
registry:
image: "gitea.dooplex.hu/admin/felhom-controller"
username: "" # Gitea credentials (for version checker + template fetcher)
token: ""
check_interval: "6h"
template_interval: "1h" # How often to refresh controller.yaml.example
Files added
internal/configgen/configgen.go— shared YAML generation packageinternal/web/configs.go— web handlers for config CRUDinternal/web/templatefetcher.go— background template refreshinternal/web/controller.yaml.default— embedded fallback templateinternal/web/templates/configs.html— config list pageinternal/web/templates/config_form.html— create/edit forminternal/web/templates/config_detail.html— detail + credentials page
Files modified
internal/store/store.go— customer_configs table + CRUD methodsinternal/api/handler.go— config retrieval endpoint, per-customer auth,ConfigTemplateProviderinterfaceinternal/web/server.go—/configs/*routes,SetTemplateFetcher()internal/web/embed.go— embedded default templateinternal/web/templates/dashboard.html— navigation barinternal/web/templates/customer.html— navigation barinternal/web/templates/style.css— form, nav, button, credential stylescmd/hub/main.go— template fetcher wiring,TemplateIntervalconfigconfigs/hub.yaml.example— registry section
v0.1.8 (2026-02-16)
- Controller update trigger: "Update" button on customer detail page calls controller's self-update endpoint
- Registry version checker: background goroutine checks Gitea registry for latest controller image tag
- Update available indicator on customer detail page
v0.1.7 (2026-02-15)
- Infrastructure backup endpoints for disaster recovery (POST + GET
/api/v1/infra-backup)
v0.1.6 (2026-02-14)
- Handle disabled reporting status
- Storage labels display
- Date in history table
v0.1.5 (2026-02-13)
- Notification preferences sync endpoint (
POST /api/v1/preferences) - Notification display on customer detail page
v0.1.4 (2026-02-12)
- Resend API key support for email notifications
- Notification endpoint (
POST /api/v1/notify)
v0.1.3 (2026-02-11)
- Customer detail page: system info, storage bars, container table
- 24h history graphs
v0.1.2 (2026-02-10)
- Dashboard auto-refresh (60s cycle)
- Status logic (green/yellow/red based on report age + health)
v0.1.1 (2026-02-09)
- Basic dashboard with customer overview table
- Report ingest API
v0.1.0 (2026-02-08)
- Initial release: SQLite store, report API, basic web dashboard