R-36: both halves delivered — the enabled-but-unprovisioned warning on the
customer page (reusing the same predicate the offsite re-issue handler refuses
on), and the related sub-item, auto-minting the self-bind link at customer
creation AND RESET completion so the console banner's promised email is already
true. Records the gap found while wiring it: PurgeCustomerResetDBState does not
clear selfbind_tokens, so a pre-RESET link would have survived the reset; the
skip paths now clear stale tokens.
R-37: the post-RESET staleness banner, narrow by design — an in-flight reset
does not trigger it, it clears itself on the first post-RESET report, and ties
resolve to STALE because SQLite timestamps are second-resolution and a
same-second report almost certainly predates the reset.
Both red-proofed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
Deploys hub v0.67.0 (auto-minted self-bind link, post-RESET staleness banner,
unprovisioned-offsite warning, pbsdr_reissued flash text). The manifest is the
truth — the code push and image build deploy nothing until this tag moves.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
Four small items, each a case where the hub already knew something and said
nothing. Green: build, vet, tests all pass.
(a) Self-bind link is minted automatically at customer creation AND at RESET
completion (R-36 sub-item). The console banner tells the customer to open
"az e-mailben kapott link"; until now that email existed only once the
operator remembered the button, so the banner could point at something that
did not exist — during the 2026-07-18 rehearsal the box waited ~11.7 min on
exactly that. handleSelfBindLinkSend's body was extracted into a shared
mintAndSendSelfBindLink core so the button and the auto-mint callers cannot
drift apart on the honesty rules: F1 (no address -> mint nothing) and F2
(send failed -> delete the token, never leave it live). The wrapper NEVER
fails the operation it rides on — a create that provisioned Cloudflare,
offsite and PBS must not 500 over a courtesy email.
Gap found and closed while wiring it: PurgeCustomerResetDBState does NOT
clear selfbind_tokens, so a link minted BEFORE a reset would have stayed
live across it. A successful mint already replaces it (delete-then-insert,
single-active); the skip paths would not have, so they now clear stale
tokens too. Invariant: after auto-mint runs the only live link is one it
just issued, or none.
(b) Post-RESET staleness banner (R-37). When a RESET COMPLETED after the newest
report, every health figure on the page describes a lifecycle that no longer
exists, and the page kept showing pre-RESET warnings as current. Narrow on
purpose: an in-flight reset does not trigger it, and it clears itself when a
report arrives. Ties resolve to STALE — SQLite timestamps are second-
resolution and a same-second report almost certainly predates the reset;
erring the other way would hide the banner exactly when it matters.
(c) Unprovisioned-offsite warning (R-36 interim). enabled==true with type=="" is
a real, stable, silent state: provisioning is Save-triggered and the
re-enroll auto-re-issue deliberately skips an unprovisioned target, so
nothing self-heals it. Reuses the exact predicate the offsite re-issue
handler already refuses on.
(d) pbsdr_reissued rendered an EMPTY flash box — the key had no template branch,
so re-issuing PBS credentials showed a success box with no words (observed
live 2026-07-18). Now describes what was staged plus the R-39 caveat:
confirm `pvesm status` shows the entry active, because a converged agent can
report `applied` while the storage still 401s.
New .flash-warn (amber, --warn tokens) for the deviation tier between success
and error — exception-color principle: only on deviation, never on a healthy
page.
Tests assert each banner is ABSENT in the nominal cases as well as present in
the deviating one — a banner that always renders is worse than none. Both
red-proofed: deleting the pbsdr_reissued branch reproduces the original empty
box; neutering the staleness predicate fails the banner assertion. New
read-only store accessor CountSelfBindTokens makes the single-active invariant
assertable.
NOT in this train: the R-39 hub-side generation-bump fix the pre-travel task
made conditional. Its condition was REFUTED (SetHostDesired bumps
unconditionally; applyPBSDR is idempotent as documented) — the real mechanism is
the agent's descriptor-hash convergence and needs its own spec.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
Follow-up to be2fc50 — the index "Technológiák" preview still carried a
"Kubernetes / Üzleti szintű rendelkezésre állás" tile, an availability
promise with no capability-map row, pointing at a section that commit
removed. Replaced with the map-backed two-tier backup (§C tier-2 +
offsite, both PROVEN-LIVE).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N1W4wBum4JSFrbaEoDkMBy
Audience shift: a Facebook post recruiting volunteer testers is about to send real
Hungarian households (mostly on phones) to a site that until now had zero stakes.
Every claim re-checked against documentation/architecture/00-capability-map.md.
- Naming ruling: "Felhő Felügyelő" removed site-wide (14 occurrences, now 0). The
brand is Felhom; the interface is the vezérlőpult.
- index.html: new "Mit tud a doboz ma?" (8 map-traceable cards, incl. Hálózati
megosztás and the customer-only recovery code) + new "Zárt teszt" section with
stated limitations (one shared household password; TV-re streamelés hamarosan).
CTA reuses the existing live contact-mailer via /kapcsolat?tema=zart-teszt.
- og:image was a site-wide 404 (pages pointed at a .png that never existed) —
generated a branded 1200x630 card + width/height/alt. Load-bearing for the post.
- App count 45+ -> 53 (real catalog count).
- Cut unbacked claims: the Kubernetes/k3s section + multi-node tier, Tailscale ->
WireGuard, the RAID card -> honest two-tier backup, gyik multi-user answer
(both JSON-LD and visible copies), and the "azonnal értesítést kapsz" overclaim.
site_gates.py green. Mobile measured at 380px: scrollWidth == clientWidth == 365.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N1W4wBum4JSFrbaEoDkMBy
Golden 0.146.0 baked on the drill VM and published to gitea:
felhom-golden/0.146.0/golden.tar.zst
sha256 4834c703162c5437467a329144b1a523019bf5693ab9d439558be7323587e955
612696588 B (584 MB archive), controller 0.146.0 confirmed baked in
All pass markers green: Result=success/ExecMainStatus=0, 0 FATAL/exclusions,
docker OK (overlay2), ALL THREE mounts included (rootfs + mp0 /var/lib/docker +
mp1 /mnt/sys_drive), pre-delete HTTP 404 (the pre-gate — version did not exist),
upload HTTP 201.
Integrity verified independently of the build host: anonymous GET | sha256sum
matches byte-for-byte, ranged GET 206, content-length matches the bake's bytes.
The version now appears in the hub dropdown (0.136.0, 0.143.0, 0.146.0).
Teardown per GL-1: log copied out as evidence first
(180:/mnt/5_hdd/felhom.eu/drill/bake-0.146.0.log), guest 9100 purged, token +
script + log shredded in-VM, VM off, qemu confirmed gone via `ps -eo comm` (not
the self-matching pgrep -f), drill disk reverted to the virgin snapshot exactly
as found. Token-leak grep = 0 against the LITERAL token value, on the bake log
and both ISO build logs from this session.
REMAINING is operator-only and password-gated: Day-0 manifest Golden -> 0.146.0
(Agent stays 0.90.0, MinAgent stays 0.90.0 — v0.146.0 declares no new agent
coupling), then the floor -> v0.146.0 saved LAST.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
ROADMAP:
- R-39 gains the full live diagnosis and REFUTES the brief's hypothesis. The
generation IS bumped (SetHostDesired bumps unconditionally, 2->3) and
applyPBSDR is exonerated, so no hub fix was shipped. The real mechanism is a
signal mismatch: the hub's re-consume signal is a generation bump + poke,
while the agent re-applies on a change of the DESCRIPTOR CONTENT HASH
(manager.go ~L235). An ep0 re-issue re-keys the secret of an EXISTING token,
so token_id/fingerprint are unchanged, the descriptor is byte-identical, the
hash never moves, and the fresh secret is never consumed -> 401 forever.
Proof: consumed-failed.json carries the same hash a4e5424... as the marker
written two minutes before the re-issue.
Records the second defect found while healing (wrapper reconcile passing
--server, fixed in agent v0.90.1), marks the box HEALED with evidence
(pvesm active, token 200, a real 9.7 GB encrypted backup listed PBS-side),
and leaves the fleet fix explicitly pending its own spec.
- R-33 collapses to SHIPPED (scripts v1.21.0), incl. why
TimeoutStartSec=infinity is the load-bearing half.
- Pre-invite checklist: golden target moves 0.145.x -> 0.146.0 and notes it is
now MORE stale, since v0.146.0 is live on the demo box while the golden still
bakes 0.143.0.
REPORT overwritten with the train: R-39 diagnosis verbatim + heal evidence, the
two ISO shas with the byte-identical-payload verification, the nav polish and
why the screenshot leg could not be done (the demo controller password is
customer-owned since the claim flow, so the build-server credentials are stale),
Phase 4 skipped cleanly, and Phase 5 deferred rather than half-run.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
Waiting to be bound is the NORMAL state of a freshly installed box, and it must
not be reported as failure. The PAIRING poll loop used to BE systemd's
Restart=on-failure/RestartSec=30 — one poll per invocation, exiting non-zero
until the bind landed — so every 30s systemd printed "Failed to start Felhom
host bootstrap" on the physical console the CUSTOMER is watching. The
2026-07-18 N100 rehearsal measured 52 FAILED lines in ~11 minutes while nothing
was wrong (VALIDATION-n100-rehearsal-2026-07-18.md F6).
felhom-bootstrap.sh: run_pairing() is now a while-loop that sleeps
POLL_INTERVAL (30s — the hub-side rate is unchanged) between polls, so the unit
sits in `activating`. Registration split into register_appliance(), which
returns non-zero for a transient problem (no network yet, no identity, no
token) and is retried by the loop instead of taking the unit down. Cadence
constants: POLL_INTERVAL=30, BANNER_EVERY=10 (5 min), HEARTBEAT_EVERY=20
(10 min).
Quiet without going dark: a 204 is logged once on entry (worded so nobody reads
it as an error) and then only on the 10-minute heartbeat with elapsed minutes;
404 and unexpected codes degrade the same way. 410 STILL exits non-zero on
purpose — delivery consumed but no local env is a real crash window, and a
clean systemd restart is the right response.
Console banner: every 5 min instead of every cycle, single accented spelling
instead of the parositasra/párosításra double, and the reassurance the
rehearsal showed was missing ("Ez a képernyő magától frissül — nincs teendő a
doboznál").
felhom-bootstrap.service: TimeoutStartSec=infinity. This is load-bearing, not
cosmetic — a Type=oneshot ExecStart is killed at DefaultTimeoutStartSec (90s),
so without it systemd would kill the new in-script wait after 90 seconds and
Restart=on-failure would silently reinstate the exact spam this removes, after
appearing to work for the first three polls. Restart=/RestartSec= are kept
deliberately: they still cover the DIRECT path, a failed host-install, and 410.
Verified behaviourally, not assumed: driven in a throwaway Debian container
against a stub hub answering 204 five times then delivering — logged the wait
once plus one heartbeat, never exited between polls, then consumed the
delivery, wrote the 0600 env, fell through to the direct install in the same
invocation and exited 0. The old design produced five unit invocations and five
"Failed to start" console lines for that same sequence.
Hub endpoints, payloads, polling rate and one-shot delivery semantics are all
unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
Overwrites REPORT.md per convention: evidence bundle manifest, the map rows
flipped with citations, ROADMAP IDs assigned (R-30..R-39 + R-27c), the seven
discrepancies found against the brief, and the remaining-to-first-invite line.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
The 2026-07-18 N100 rehearsal ran the complete final-product flow on real metal
in one pass: RESET -> generic pairing ISO v1.20.0 -> customer self-bind -> day-0
-> managed-floor lift -> escrow ceremony -> offsite snapshots. No code changed;
every finding is recorded and ranked, none fixed.
VALIDATION-n100-rehearsal-2026-07-18.md — run context, a UTC-normalised timeline
built from the hub events stream / hub DB / controller log / bootstrap + agent
journals, per-ledger verdicts for S1-S8 + ledgers 8 and 9, 12 findings, the
not-exercised list, and 7 discrepancies against the brief.
Headline wall-clocks: bind -> credential 26 s; bind -> controller running the
current version 2 min 44 s; managed floor 0.143.0 -> 0.145.0 in 5 s unattended
(initiated_by: auto-floor); escrow ceremony -> offsite enabled 12 s; drive enrol
30.3 s. No post-bind leg stalled, which is the immediacy row's real-onboarding
proof.
Capability map (10 citations added):
- Bare-metal Felhom ISO PARTIAL -> PROVEN-LIVE (F1 closed on metal)
- Customer self-bind (slice 1) IMPLEMENTED -> PROVEN-LIVE (customer_selfbind)
- Guest RAM resize (R-24) IMPLEMENTED -> PROVEN-LIVE (shrink AND grow)
- Customer RESET two real firings + verified external teardown
- Escrow ceremony first live wizard firing
- Immediacy row "real-onboarding proof pending" cleared
- Publish train box-side floor lift proven on a fresh install
- Customer claim R-4 gmail half (Inbox under p=quarantine)
- Offsite orphan guard staged live leg fired on its own
- DR tier by default candidate PROVEN-LIVE upgrade WITHDRAWN (R-39)
Not flipped, as instructed: customer-performs-restore, BYO, DLNA, multi-user.
ROADMAP — collapsed R-1 (appliance half done, Peti half survives), R-21
(physically closed), R-24, R-27 slice 1, R-4. New ranked items:
P2-HIGH R-39 PBS DR applied-but-dead R-30 liveness from the wait channel
R-31 async offsite + status R-32 RESET base-dir purge
R-33 bootstrap quiet-poll
P2 R-34 backup lifecycle R-35 config-apply session survival
R-36 post-RESET offsite prompt R-27c console-passphrase bind
P3 R-37 post-RESET health card R-38 installer GRUB slice
Plus a pre-invite checklist (golden 0.145.x rebuild, freemail.hu, C6, R-11).
R-39 is NEW and was not on the brief: the PBS DR descriptor auto-provisions and
the agent converges state=applied, but pvesm reports 401 Unauthorized/inactive
and a direct probe 401s on every endpoint including /version while WG is healthy.
The hub minted a second token secret two minutes after the agent applied the
first and consumed_at is still NULL; the converged state machine will not
re-apply, and the agent's verify loop cannot read the credential to notice it
(non-root read of a file it writes through a root wrapper). Rank is provisional
pending Viktor.
R-3 draft: all four [REFINE] slots filled, self-bind made the default path with
"send the link BEFORE the customer sees the console", the measured wall-clock
table added, and interim operator workarounds for R-31/R-36/R-39. C6 (renumbered
C7) is marked as the single unexecuted step and keeps the doc a DRAFT.
Evidence bundle: 180:~/n100-rehearsal/ (10 files + MANIFEST.md), collected before
the box was unplugged for travel. Secrets read only to run probes; recorded as
lengths and metadata, never values.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
Origin: R-7b close-out (felhom-controller REPORT section 4f). Two parts:
(a) the finding is benign — estimate.go:179 mounts a NAMED VOLUME (daemon-side, no
host path), the same shape as three already-allowlisted entries, so the fix is a
3-line ALLOWLIST addition with its WHY, NOT a docker-cp rewrite;
(b) the systemic half: the gates run only when a human remembers, so this one sat
red from v0.129.0 (2026-07-14) to v0.145.0 while REPORTs said green. Second
instance of the class after the v0.123.0 'Windows green gate silently red' note.
Lists the full gate inventory to audit for the same rot.
Drops the false 'no offsite target on the demo box' clause from the ROADMAP row,
the capability-map SMB row and sharing.md. Cites offsite snapshots e0b9d723 /
4e2b15ec and the restore round-trip results. Root cause (guessed settings key) is
recorded in felhom-controller REPORT section 7b.
- capability map: SMB row KNOWN GAP cleared -> share data rides both tiers; the
offsite leg + restore round-trip flagged as not-yet-live-exercised
- ROADMAP R-7b: idea -> SHIPPED, with the Model B' rationale and the live evidence
- controller/sharing.md: the KNOWN GAP block replaced by the execution contract;
operator note corrected — samba IS liveness-monitored since v0.145.0
Viktor's human leg closed the last gate: both shares open from the Windows
Network view, an interactive Explorer save landed as uid 1000, and a write into
the read-only share was refused with the folder untouched. Capability map row
flipped to PROVEN-LIVE with that evidence; ROADMAP R-7 + sharing.md updated.
R-7b (shares classified but not in any live backup run) remains open.
controller/sharing.md (code-verified vs controller v0.144.0 + felhom-samba
1.0.0); capability map 'Files from Windows Explorer / Mac Finder (SMB server)'
MISSING -> IMPLEMENTED (PROVEN-LIVE pending Viktor's Explorer leg); ROADMAP R-7
-> shipped-slice-1 with the slice-2 remainder, and the backup design fork split
out as R-7b (shares are classified but not in any live backup run yet).
Human Explorer leg exposed the split: FELHOM-SPIKE renders (WSD PASS) but the
double-click fails 0x80070035 — flat name resolves by no path (DNS/LLMNR/NetBIOS
all silent; disable netbios=yes). By-IP mount works => SMB is healthy, the gap is
name resolution. Fix verified live: adding nmbd (NetBIOS) => nbtstat lists
FELHOM-SPIKE, ping resolves, \FELHOM-SPIKE\spike-share mounts by name. R-7 must
ship smbd+wsdd+nmbd (+avahi/.local), not wsdd alone.
Verdict: appliance guest is LAN-bridged; multicast discovery works only in the
guest netns (guest-direct or docker --network host) — default bridge is deaf to
LAN multicast. Real samba+wsdd on host-net: Windows 11 ProbeMatch + 445 + SMB
round-trip PASS; SSDP MediaServer:1 reaches LAN clients. R-7 => host-network
LAN-bound; R-8 Jellyfin-DLNA plausible if host-network. ROADMAP R-6 -> spiked,
R-7/R-8 unblocked. S4.4 Explorer render pending human.
Live through the real ingress: public /bind/ renders logged-out with the
no-oracle expired state (200-not-500 proves selfbind_tokens migrated);
gate intact (/ and /hosts -> /login); CSRF exemption is /bind/-only
(POST /bind/ no-CSRF 200 vs POST /customers/x/block no-CSRF 302).
Operator-minted full walk + new-ISO console banner remain operator/R-1.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017qDiBqKKQ5vPB5fXBqu7Kp
Let a customer bind their own freshly-installed appliance without the
operator: operator "Send self-bind link" mints a 7-day tokenized
capability link, emailed (Hungarian, sibling sender) to the customer, who
opens a public /bind/<token> page and proves two factors — the console
pairing code shown on the box screen + their retrieval passphrase — and
the hub stages the bind via the same BindAppliance (provenance
customer_selfbind). The box's ~30s appliance poll delivers.
Viktor's three rulings verbatim: console pairing code (no appliance list
ever rendered), operator-sent tokenized link, 5-attempt lockout ->
"call support". Wrong code == wrong passphrase (one generic failure, no
oracle, both factors compared unconditionally); expiry falls back to
operator-bind unchanged.
THE TRAP: one public prefix /bind/, exempt from auth+CSRF at both /login
gate sites via a single isPublicBindPath predicate (tight trailing-slash
match; ServeMux ..-cleans; handler rejects '/' in token). 9 tests
(Scenarios A-F + F1/F2); 4 red-proofs verified red-then-green (lockout,
oracle, widened-prefix, single-active). GC verdict: no appliance GC ->
the 7-day TTL stands alone. Controller/agent untouched; R-27b deferred.
Green: full hub build/vet/test (17 ok) + bash -n + hub confirm gate.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017qDiBqKKQ5vPB5fXBqu7Kp
Makes PBS DR storage visible like the restic pool box (v0.64.0), differentiated. Scoping
correction: restic = subaccounts on the shared Hetzner Storage Box (Hetzner API); PBS DR =
the felhom-offsite PBS datastore on the ep0 endpoint VM (NO Hetzner API). Option A
(Viktor-ruled): a read-only `usage` op on the felhom-tenantsync ep0 forced command (twin of
fingerprint), polled by a new hub checker on the 15-min throttle. READ-ONLY throughout.
Phase-0 (gate PASSED): on ep0 (PBS 4.2.3), df -B1 --output=size,used,avail <datastore path>
yields bytes (39990112256/7627939840/... ~19%), read-only, existing sudo context, no admin token.
- scripts/felhom-tenantsync.sh -> v1.2.0: read-only `usage` short-circuit (df on the datastore
path), no customer_id, no admin token, NO mutation. + a bash harness proving zero mutation.
- tenantsync.Client.Usage() + BoxUsage; unknown-op -> typed ErrUsageUnsupported (graceful).
- monitor.PBSDRBoxChecker: OffsiteBoxChecker clone over a usageReader seam; 15-min throttle,
cached PBSBoxSnapshot, escalation-only pbsdr_box_fill on the "pbsdr-box" scope (operator only,
no SaveEvent), recovery re-arm. Fill only. THREE states: ok / unavailable (ep0 <=v1.1.0,
neutral no-alert) / degraded (exec failed, keep last).
- config: Alerting.PBSDRBoxFill{Warn,Crit}Percent (80/90); built with the tenantsync client,
60s sweep, SetPBSDRBox. Hub deploy INDEPENDENT of the ep0 update (graceful degradation).
- web: /offsite splits into Restic + PBS DR hash tabs (endpoint cards under PBS DR); PBS panel;
the single dashboard tile becomes two gauges (RESTIC pct.ratio, PBS DR pct / n/a).
- runbook offsite-endpoint.md 10: v1.2.0 update steps (no sudoers/authorized_keys change).
Tests: 10 Go + the harness; 3 red-proofs (usage mutation, escalation-only, unavailable-drives-band)
confirmed red then restored. go build/vet/test + bash -n + hub confirm gate all pass.
The operator sees the shared pool box's real state on the hub: total box fill vs
capacity, Σ(shared soft quotas) vs capacity (the oversubscription ratio), per-customer
usage/quota bars, and a box-level operator alert (fill % + oversub ratio) on the existing
dispatcher's operator channel. Per-customer fill alerts already existed; the box-level
aggregate was the gap. READ-ONLY against Hetzner (GET only).
Phase-0 probe (gate PASSED): the live pool box 611714 returns capacity via
storage_box_type.size (1 TiB / bx11) and usage via a stats object (size/size_data/
size_snapshots), all bytes; our token reads it (200).
- hetznerapi: additive StorageBoxType + StorageBoxStats on StorageBox (no existing field/
method changed); fake carries them + a GetBoxCalls counter; golden decode test.
- monitor.OffsiteBoxChecker: OffsiteChecker-sibling for the box; fetch-throttled (1 GET/
15min), cached BoxSnapshot, escalation-only + recovery re-arm. FILL (used/capacity 80/90)
+ OVERSUB (Σ shared+enabled quotas / capacity, 2.0x) — independent. Σ from the ConfigJSON
Descriptor (offsite.ReadDescriptor, new), never the report echo; dedicated+disabled
excluded. Scope "pool-box" -> operator channel only, no SaveEvent. Failed fetch keeps the
last snapshot degraded; missing data never becomes 0% and never transitions a band.
- config: Alerting.OffsiteBoxFill{Warn,Crit}Percent + OffsiteOversubWarnRatio (80/90/2.0
defaults; thresholds pending Viktor's ruling). Constructed in the HETZNER_TOKEN branch,
60s sweep, snapshot handed to the web server.
- web: Offsite-tab panel (fill bar, Σ+ratio, per-customer usage/quota rows) + a compact
dashboard tile; reads the cached snapshot only, never fetches; nil -> "not configured".
Tests: 10 new + 4 red-proofs (throttle, Σ filter, escalation-only, failed-fetch honesty),
all confirmed red then restored. go build/vet/test all pass; hub confirm gate OK.
The immediate-sync arc covered only operator-initiated desired-state changes;
system-initiated mutations bumped the generation silently, so a freshly onboarded
box waited a full agent tick for state the hub had already minted (observed live at
slice-C onboarding). Wire the existing, live-proven notifiers into every system site
on the correct plane — call-site wiring only, no new mechanism.
Agent plane (poke.Notifier):
- web/pbsdr.go: PBSDRAutoProvision (the observed lag), ReissuePBSDR (also lifts the
pbsdrheal reconciler escalation, zero reconciler changes), handlePBSDRReissue —
each pokes AFTER the successful SetHostDesired, never on a blocked/error path.
- api: new nil-safe Poker seam (PokeHost/PokeAllHosts + SetPoker); handleAdminSetDesiredState
pokes the target host; handleAdminSetOperatorPeer fires PokeAllHosts only when the
fleet generation bump succeeded (fire-after-commit).
- main.go: one poke.Notifier now feeds both planes (SetPoke + SetPoker).
Controller plane (intent.Hub.Bump):
- api/reissueOnReenroll: one nil-guarded bump so a long-polling controller wakes in
seconds instead of on the 15-min cycle.
Deliberate non-sites (unchanged): WG register (undeliverable pre-tunnel — the agent
fast-tick SECONDARY owns it), WG delete (transport removed), pbsdrheal Restage (no
generation bump → the 60s ticker is the pickup path). internal/pbsdrheal byte-unchanged.
Tests: 10 non-hollow tests (web async channel-synchronized fake sender; api synchronous
fake Poker) with explicit zero-count negatives; representative red-proofs per group
(A/B/C/D) run-fail-restored. Green: go build/vet/test all pass.
A generic ISO carries NO customer secret. The box registers itself at the hub
as an unclaimed appliance; the operator binds it to a customer; the hub delivers
the customer-id + retrieval passphrase ONCE; day-0 completes via the slice-A path.
Hub (v0.62.0):
- store/appliance.go: appliance_registrations keyed by (uuid, mac_set) — MAC set
is the tiebreaker (duplicate SMBIOS UUIDs); token stored as sha256 only.
Idempotent register (sticky-discard), atomic one-shot delivery, bind/discard.
- api/appliance.go: POST /appliance/register (the one unauth endpoint, per-IP
rate-limited, 256-bit token); GET /appliance/poll (404 no-oracle / 204 unbound
/ 200 deliver-once / 410 delivered). Passphrase read live, never logged.
- web/appliances.go: Hosts-page "Unclaimed appliances" section + BIND (customer
picker, host count display-only) + DISCARD; SSH host-key fingerprints; events.
- Red-proofs: one-shot delivery + register idempotency (both proven red);
404-no-oracle, sticky-discard, bind staging, render. Green + confirm gate.
Scripts (v1.19.0):
- felhom-bootstrap.sh: ONE unit, TWO modes. Direct (env has customer/passphrase)
= slice-A path, byte-identical, only branched around. Pairing (generic) =
register + poll (RestartSec=30 is the poll timer); on delivery write the env
0600 and fall through to direct. Secrets + token shredded on success.
- build-felhom-iso.sh --pairing: generic secret-free ISO, -generic filename,
manifest mode=pairing. profiles/generic.profile (new).
- test/bootstrap-modes.sh: Scenario D (direct = zero appliance calls) + pairing
register/poll + delivery handoff — all green in a debian container.
Closes N100 F1 (HIGH): cheap AMI (AN3PLUS 0.01-class) UEFI firmware can't
relocate the ISO's stock signed GRUB from USB (relocation 0x0). The run's live
grub-mkimage workaround is now a first-class pipeline mode.
- build-felhom-iso.sh: --loader shim|mkimage (default shim, byte-for-byte
unchanged; profile-settable FELHOM_LOADER; --loader wins). Loud banner +
manifest loader:/grub-mkimage: fields + -mkimage filename suffix.
- mkimage-surgery.sh (new): post-prepare-iso, in the assistant container. Builds
a monolithic grub-mkimage loader from the ISO's own GRUB (module set from its
grub.cfg; embedded search --fs-uuid -> configfile the real menu). Swaps it into
the ISO9660 tree (real lowercase path) + the efi.img ESP; xorriso re-master
preserves BIOS-hybrid + UEFI + GPT-ESP, drops only Apple HFS+/APM. Recipe from
the N100 run evidence, not re-derived.
- Dockerfile.assistant: grub-common + grub-efi-amd64-bin + mtools + dosfstools.
profiles/n100.profile (new, mkimage + SB-off note).
- Validated on nested VM 311 (RUNBOOK-B legs): leg1 shim boots+installs under
OVMF SB-enforcing + SeaBIOS; leg2 mkimage boots+installs under SB-off; leg3
(red-proof) mkimage under SB-enforcing FAILS Access Denied (unsigned -> SB must
be OFF); leg4 surgery byte-identical payload. bash -n + shellcheck clean.
Physical N100 closure folds into the rehearsal (n100-safety match-nothing ISO
built + sha-recorded, unbooted). PXE stays a deferred R-21 note.