Part A evidence (migration spike, sftp-written repo through the pinned rclone key) and the hub
red-proofs under documentation/audits/offsite-lock-build-2026-10-03/. Manifest bump follows after
the image is built and Secret/offsite-secret-key exists.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Allowlisted, not operator-only, seeded for new households and added
once (add-only) to every existing enabled_events row. mail.event entries
in hu and en name the app from details.stack_name. Per-app cooldown on
both the operator and the household leg.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found during v0.118.0's own live proof, which is the only reason it was found:
I went to press "send test notification" for the English demo box and read
sendTestEmail first.
It had its own hardcoded Hungarian subject and body and never went through
FormatCustomerEmail, so it was the one customer mail v0.118.0 did not localise
— and it is the only customer mail an operator can trigger on demand, which
makes it the one most likely to be used to check whether the localisation
works. Pressing the button for an English household would have answered that
question wrongly, and convincingly.
The two sentences are extracted byte-for-byte into the bundle, so the Hungarian
test mail is unchanged. Red-proofed against the hardcoded version.
The general form worth keeping: the surface you would use to CHECK a feature is
the one most worth checking first. A broken instrument that reports success is
worse than a broken feature.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The hub has written every customer e-mail in Hungarian whatever the box was set
to. The box has published its language since controller v0.247.0; nothing read
it. Now it does.
Nothing an operator reads changes. The Hungarian mails are byte-identical, and
that is a diff rather than a reading: 56 goldens per language captured from
v0.117.0 BEFORE any string moved, and all 56 Hungarian ones pass unchanged after
every sentence was routed through the new bundle.
- internal/i18n: flat bundle, 79 keys, hu authoritative + hu fallback, ceiling 0.
- customerMessages/severityLabels are DERIVED from the bundle, so a sentence is
written in one place and all 40+ tests that read those maps still work.
- Language order: last reported -> created-with -> hu. reports.language defaults
to EMPTY, never hu: "never told us" is not "chose Hungarian".
- message_customer on POST /api/v1/event, additive and optional forever, for the
sentences the box composes and the hub cannot translate.
- The bind page is per-language, and its `expired` state stays Hungarian: it is
the state an unknown token lands in, so rendering a real English customer's
token in English would make the LANGUAGE answer what the TEXT refuses to.
Two defects found inside the release:
- R-581: the newest report was picked by received_at, which has SECOND
granularity, so same-second reports tied and the winner was arbitrary. Ordered
by the autoincrement id now. GetCustomers() still has the shape - row open.
- R-582: the English copy-guard stems, ported word for word from Hungarian,
convicted 141 honest sentences. The English claim is a phrase with a modal.
R-555 closed: the language allowlist entry is out of wire_contract_gate.py.
hub_copy_gate.py follows the sentences into the bundle - without that it would
have scanned four files that no longer hold any customer text and reported
success. Three new decoys incl. an innocent control.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-549 (operator ruling A): controllerStatus hardcoded 30m/1h while both
staleness checkers and hostStatus read alerting.stale_threshold. Moving the
threshold to 45m would have painted a customer amber 15 minutes before the
alarm could fire - the second definition rollup.go's header forbids. It now
reads the same value, down at 2x. Both 'checker initialized' log lines print
the threshold, which no line did before.
R-539 (ruling 3 of 2026-09-16): controller_slow_crashloop (warning,
operator-only), minted when the agent's slow_crashloop_since moves, with the
fast sibling's first-sight rule.
R-550: restore_interrupted (warning, for the household) allowlisted with a
Hungarian customer message.
Red-proofs, each seen failing then passing: the status test with the old
hardcoded numbers; the checker test with the movement branch removed; the
operator-only test with the registration removed; the household-message test
with the Hungarian entry removed (asserted on the SUBJECT - the body
legitimately repeats the raw message, which my first version of the test
mistook for a fallback).
go build/vet/test ./... green, 18 packages.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Off-site is ON by default for a new customer — shared, 100 GB soft quota prefilled,
the checkbox kept so an operator can opt a customer out. The reason is this repo's
own [FACT]: the whole-guest tiers do not carry the data drive and a Tier-1 unit has
no file leg, so with this unticked a one-drive box keeps NO copy of the household's
own files. Measured on a fresh box the same day.
The quota is prefilled because the fill warning only fires when quota_gb > 0.
Also registers controller v0.244.0's app_deploy_started / app_deploy_failed in both
allowedEventTypes and customerMessages, per the rule that the two move together.
Red-proofed: dropping the default fails the new render test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Phase 0: the public ISO never auto-installs by construction (no answer.toml,
G1); the operator re-affirmed the interactive installer 2026-09-14.
- felhom-bootstrap.sh: mask pvebanner.service, write a Hungarian /etc/issue
(no :8006 admin URL); pairing banner names the Tulajdonosi jelmondat and
paints through the CONSOLE_DEV seam (R-496). Harness: 8 checks, red first;
fake hub now sends a pairing code (the banner was never tested, R-502).
- hub: created flash + Credentials block tell the operator to hand the phrase
over; the self-bind mail names the operator (R-497). Tests red first.
- iso-release-gate G14-G16; domain ruling in 01-topology + CONTEXT; R-494
narrowed to P3; R-502..R-504 filed; volunteer guide and day-0 A.2 aligned.
ISO_VERSION 1.27.0 (not built, not published).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
THE RECORD WAS TELLING A WORSE STORY THAN THE TRUTH FOR TWO MONTHS, and my own probe is why.
R-429 CORRECTED. Yesterday's spike searched for a directory called `.snapshots`. The vendor documents
the path as /.zfs/snapshot. The probe's CONTROLS were sound and its SUBJECT was wrong, so "not found"
was true and meant nothing. Re-probed at the documented path on both boxes, with controls:
- /.zfs lists (shares, snapshot) from inside the jail;
- a write into /.zfs/snapshot is REFUSED - dest open ...: Failure - while the identical write to
the account home SUCCEEDS and was cleaned up.
That is the append-only property PROVEN rather than cited, and it is the sentence the whole re-scope
rests on. Seven daily snapshots are confirmed in the panel. The mitigation works. What remains true,
and was always the actual finding: the row claiming it had no R-number, its "confirm tomorrow" went
36 days unanswered, and the DUE-CHECKS block built for that class was empty. THE FINDING WAS NEVER THE
SNAPSHOTS - IT WAS THAT NOBODY COULD TELL.
R-95 RE-SCOPED: the box can delete its LIVE repository but cannot write to the daily snapshots of it,
so a deletion costs at most one day plus a per-file recovery - not open-ended loss. The ranking is
Viktor's; it has been #1 since July on the old story.
R-432 FILED: a sub-account sees /.zfs/snapshot EMPTY while the same box holds seven snapshots, so
per-file recovery is operator-only today. One panel read settles whether a NAMED snapshot can still be
entered, which would make it product-reachable.
R-431 SHIPPED. Third signal in OffsiteChecker. On the hub deliberately: a detector on the box is one
the deletion can silence. Threshold REASONED, not invented - over 12 898 reports every decrease lands
on ZERO and predates stats_known, and in the stats_known window there are none, so observed churn gave
nothing to calibrate against. Retention cannot halve a total; a mass deletion goes to ~0. Hence: more
than half, and at least 5. Guarded by StatsKnown (R-331), the declared State (R-204) and run success
(R-100 - whose lesson lives in this very file).
ACCEPTANCE: 9 009 real report points replayed through the detector produced ZERO alarms.
Three red-proofs run. The escalation one only became real after the first version was found HOLLOW -
it re-swept the same report, so the baseline had already moved and the latch was never consulted.
07 row 10's status is NOT moved: the write-refusal is measured, but the recovery ROUTE has never been
walked, which is what PARTIAL means.
Controller v0.231.0 adds a nightly job that proves an app's newest off-site snapshot still
CONTAINS that app's data. When it finds one that does not, it emits offsite_proof_empty.
Two register lines, both load-bearing and both in this commit:
- allowedEventTypes - an unallowlisted type is answered 400 and VANISHES, so this entry is
what makes the alarm exist at all.
- operatorOnlyEvents - a missing customerMessages entry is NOT a routing block
(FormatCustomerEmail falls back to the raw message, the v0.78.0 defect that register was
built for). A customer can take no action on a hollow recovery unit.
DELIBERATELY NOT reusing backup_integrity_failed, which is the nearest existing type: it
means THE STORE IS DAMAGED and carries the Hungarian template saying so. Here the store is
sound and the CONTENT is absent - different cause, different action, and telling a customer
their backups are damaged when they are not is the more expensive mistake. Same asymmetry
looksLikeRepositoryDamage is shaped around.
DELIBERATELY no customerMessages entry (the controller's dynamic Hungarian names the app and
what is missing; a template would discard it) and DELIBERATELY not in perAppCooldownEvents
(a fenced act - this job proves ONE app per night, so the coarse hourly cooldown is already
the right grain).
This widens the R-87 task's stated repo scope to felhom.eu/hub/. The reason is recorded in
felhom-controller/CONTEXT.md ruling 4 rather than left as an unexplained diff.
Hub green gate: go build/vet/test all pass, 18 packages.
The cooldown key was customerID:eventType plus the tier and run suffixes, and
none of them names an app, so every app going down inside the same hour
collapsed onto one key and only the first was mailed. Measured on demo-hp:
bookstack sent 09:27:51, privatebin suppressed 09:31:51 under
key=demo-hp:app_start_failed.
cooldownStackSuffix is the third sibling of cooldownTierSuffix and
cooldownRunSuffix, and separate for the reason the second one's docstring
already gives: the existing two keep byte-identical semantics for every type
that uses them.
It is ALLOW-LISTED to app_start_failed and takes the event type as well as the
details, unlike its siblings, and that asymmetry is the safety property. The
backup family's cooldown is coarse ON PURPOSE (R-97a, R-182) so one full disk
sends one digest rather than one mail per app - and crossdrive_failed is
severity error, reaches the operator leg, and carries stack_name through a
DIFFERENT struct, so a payload-shape rule would have split it silently. The
hour itself does not change.
Gate 11 refuses a push whose REPORT.md carries an observation with neither
`FILED: R-NNN` nor `NOT-A-FINDING: <reason>`. It deliberately does NOT accept a
passing mention of some other R-number: the lost item cited R-182 as an analogy,
so "cites a register row" would have passed the very item the gate exists to
catch. That discrepancy with the spec is recorded in the gate's docstring.
Registered here and in the controller and agent runners. NOT in the catalog
runner - it has no shared-gate mechanism and appends --all to every gate;
filed as R-391 rather than left as a sentence, which is this session's lesson.
PROMPT-TEMPLATE.md §15.9 corrected: "documented, NOT acted on" was the wording
that invited the gap, and it now names the markers and points at the gate.
R-390 filed for the golden-bake runbook's missing `pveam update`.
Hub tests 709 -> 716.
One handler, two fields, opposite discipline. An unknown event_type is rejected
with a loud 400. An unknown severity was rewritten to "info" without a word -
and severityNotifies drops "info" before BOTH legs, so the event was stored,
answered 200, and mailed to nobody.
Two shipped features went out that way: DiskAlertKind.Severity emitted "warn"
until controller v0.215.0, app_start_failed until v0.223.0. Measured on the live
hub DB today: 91 app_start_failed events stored all-time, ZERO notification_log
rows before this session - not one, on any channel.
The mechanism built to catch this class was structurally blind to it: the
dispatcher's `unrecognized severity` line cannot execute for anything arriving
over the API, because the coercion one line earlier guarantees the value it
looks for cannot arrive.
The coercion STAYS - a rejected event is a lost event, and losing an alarm is
worse than mis-routing one. Only the silence is fixed: a WARN naming the
customer, the event type and the rejected value.
The dispatcher branch is KEPT, not deleted as dead, and the reason is evidence
rather than caution: cmd/hub/main.go wires dispatcher.ProcessEvent DIRECTLY as
the monitor.EventNotifyFunc for the staleness, host-staleness and offsite-box
checkers, which never pass through the handler. For those it is the only
severity guard there is. All 90 severity literals in internal/monitor are
already valid, so the guard is silent because the producers are correct.
Test count 702 -> 709. Red-proof seen failing: delete the WARN line and the
coercion test fails with "the hub rewrote a severity and said nothing".
Golden 0.223.0 baked and published (sha 9eaf39ac3921...), round-trip HTTP 206.
Vouching is the operator's act and was not done here.
THE GAP, measured not supposed. On 2026-08-18 ep0's PBS proxy was wedged for
9 h 37 m and the hub emitted NOTHING on the operator channel. Both box
checkers hold their last snapshot and return silently on a failed fetch --
correct for a FILL signal, since a missing reading must never be read as 0%,
but it makes a dead off-site endpoint and a healthy one indistinguishable.
The only mails that morning came from the boxes' own backup failures, and
only because the WEEKLY offsite run happened to land inside the window. Two
days earlier nothing would have fired at all.
REACHABILITY is now a second, independent signal on both checkers:
consecutive failed fetch windows, reported past a default 3 windows
(~30-45 min) as pbsdr_box_unreachable / offsite_box_unreachable (warning) on
the customer-less pbsdr-box / pool-box scopes, each with a paired *_recovered
all-clear. Tunable via alerting.box_unreachable_windows (0/invalid -> 3).
THE FILL LOGIC IS UNTOUCHED. No threshold, throttle, band or escalate-once
behaviour changed; a degraded read still drives no transition.
Three decisions a later reader would otherwise "fix" back, so each is
argued in-code:
- the unreachable event REPEATS rather than escalating once. The band shape
would give exactly ONE mail at ~minute 30 of a nine-hour outage, and one
mail is missable. It leans on the dispatcher's 1 h operator cooldown to
become an hourly "still blind" heartbeat.
- ErrUsageUnsupported is NOT blindness: an old ep0 answers "no such op",
which means we reached it. Counting it would alert for days on a healthy
pre-update endpoint.
- born-blind is reported: the counter is not gated on having a snapshot, so
a hub restarted INTO an outage still speaks. last_ok is OMITTED rather
than zero-valued -- a fabricated timestamp reads as "it was fine until
then".
Both recoveries are severity "info" and severityNotifies drops "info", so
they are registered in recoveredPairedDownTypes or the operator hears that
the tier broke and never that it healed. A cross-package test drives
ProcessEvent and asserts an actual operator MAIL, not a map entry -- a green
checker test proves nothing about the seam (agent v0.91.0 shipped fully green
with SetAuthSink never called).
Tests: box_reachability_test.go (Scenarios A-F) + dispatcher_box_reachability
_test.go (wiring). Three red-proofs run and reverted, each seen failing with a
message naming the right cause: threshold 3->1, the sentinel counter guard,
the pairing entry.
Register: R-339 filed and marked SHIPPED (PROVEN-LIVE still owed -- no real or
constructed outage has exercised the emit path, and one cannot be manufactured
against Tier-2 ep0). R-340 filed: the reachability read rides ep0's LOCAL API
daemon, which the incident explicitly cleared, so this check would have shown
GREEN for all 9 h 37 m -- the honest boundary, recorded rather than glossed.
R-336's next-step corrected: pvestatd's interval is NOT tunable (Proxmox staff
have said so); the only lever is disabling the storage entry, which collides
with the agent's consume-the-one-time-secret path. Doc-only, no agent code
touched.
Hub only. No controller change, no agent change, no wire change — nothing to bake.
demo-hp untouched: the operator is re-deploying it this evening.
R-323 — the five-word phrase is „Tulajdonosi jelmondat". It was „Visszaállító
jelszó": one word from the name retired last week, and false besides — it restores
nothing, it proves the account owns the box being bound. Five sites, all in the hub;
felhom-controller and felhom-agent carry the name nowhere, so no halt and no bake.
Both suggested names were rejected with reasons: „Fiókjelszó" would collide with the
dashboard login (a DIFFERENT real secret), and „Összekötési jelszó" would leave the
two factors on this page separated only by kód-versus-jelszó — the exact shape being
removed, since the other factor is the „Párosító kód". The chosen name differs on
both axes, stem and noun. Naming only; the acceptance pin drives the real handler.
R-324 — the hub's customer copy is under a guard for the first time. Retired names
banned across all 95 hub files; retrieval stems registered in four declared customer
surfaces. The selftest found a defect in its own instrument on the first run. One
shared vocabulary in scripts/, drift-checked into the controller gate rather than
copied (R-325 removes the scaffold).
R-321 — a machine we told to be quiet is no longer reported as dead, and it was two
doors, not one: because the state is RECORDED rather than deleted, the morning
deadline check can skip it too. A deleted state returns "", which is not "down" —
R-195's shape returning through a second door. The clock runs from the report the hub
can see, so re-enabling starts it there and emits no recovery for an outage that never
happened. Three red-proofs; the one that matters showed a genuinely dead machine
sitting at "disabled" when the suppression was made unconditional.
R-326 — "which claims are unproven" is answerable by a command now. The nine I have
been repeating was the count of claims the 9 August pass DOWNGRADED, not the count of
unproven ones. The real figures: 55 claims, 23 walked, 32 not — and only 6 of those 32
cite evidence. Its first run found a stale claim (R-327).
Four paper debts and one fact given a reader. Hub-only — nothing to bake.
A4 — the entry about "the tester's machine" named a risk correctly and labelled it
in a way that invited deleting it. Established from the hub's own store: `peti-felhom`
is a REAL machine (482 reports, 2026-02-27 → 2026-07-15, a named person's own box) and
the 3.6 GB with no key and no backup is real. `david` → `tester-1` is a DIFFERENT record
with no host, no escrow and no report, ever — deleted 07:55:49 and re-created 07:56:47
this morning. The prompt's premise conflated the two; the register now says which is which.
A1 — R-312/R-313/R-303 recorded as DECIDED with their re-open triggers, and moved out
of STATUS's "Waiting on you", which is now empty.
A3 — day0-install §C.1 said pushing the installer publishes it. It has not since
R-110. Corrected, with the two manifest pins named and an outside-verification command;
the one copy that repeated it (a dated audit, true when written) carries a superseded note.
A5 — standing rule 5: evidence comes off the machine at the end of the phase that
produced it, before any revert. Earned twice in three days on the same box at the same
point (R-320). Four homes, plus what to do when it is already gone.
R-295 hub half — „Beállító kód" everywhere; „Visszaállító kód" retired. New `reenroll`
mail kind so the mail names the page a REBUILT box actually shows („A szerver
beállítása"), not the „Elfelejtett jelszó" page it has no login screen to reach.
Naming only; the acceptance pin proves the secret is untouched.
R-319 — the hub models `guest_net` after 23 days of receiving and discarding it. The
signal is `heals_last_hour`, not `state`: a guest the watchdog keeps repairing reads
healthy between repairs. `heal_succeeded` decoded too (R-260's lesson). Unknown is never
drawn as healthy — three absences, three sentences. No alarm, deliberately.
Three red-proofs, mutations asserted applied. Wire-gate checked tags 182 → 190.
B1 — the operator's 2026-08-12 dispositions were NOT in the register; they are now.
Third allowlist kind for the five ruled "no reader wanted"; `reporting_disabled`
reclassified redundant. 8 read · 5 deliberately unread · 1 redundant · 6 still owed.
Also filed: R-321 (a deliberately-silent box still alarms stale/down — the checker is
age-only, and decoding the flag would not have fixed it), R-322 (the claim guard has
never scanned the hub; a hand scan returns zero, so it is a scope gap, not a defect).
Link 6 of the recovery chain had no client. The hub has served the identity blob since
slice 10D from handleReEnroll / handleGetRestoreDirective, gated on operator-armed recovery
mode and the global key -- and nothing in the agent, the hub UI, any script or any runbook
ever called either. The only documented retrieval was sqlite3 writefile() by hand on a
kubectl cp-ed database.
GET /api/v1/hosts/{host_id}/escrow is the box-authenticated mirror of the PUT that put the
blob there. Self-scoped (a per-host key reads only its own; global may read any). A host with
no bundle gets 200 {present:false} -- a 404 is indistinguishable from an unknown host and a
bare empty 200 from a zero-length blob.
THE TRADE IS RECORDED IN THE HANDLER, not inferred: obtaining the blob used to require the
operator to arm recovery mode; now whoever controls a rebuilt box can obtain it with that
box's own credential. They still cannot open it -- the hub has never held R and a wrong code
fails closed at age's scrypt KDF. The mitigation is that every retrieval raises
escrow_blob_served (warning, operator-only), recorded before the bytes leave.
escrowSelfServiceRetrieval is the single decision point: flip it to false and the endpoint
additionally requires recovery mode, changing nothing else.
The operator-driven DR path is untouched, pinned by a test. Red-proofs observed: removing the
ownership check serves host B's blob to host A; removing the record makes it silent.
R-198 — host_escrow_superseded shipped with `blob` (the K-escrow / PBS datastore key) and
identity_blob was added to host_escrow LATER, never here. The offsite restic REPOSITORY
password lives in identity_blob. So demoteCurrentEscrowTx -- whose own comment calls it "THE
ONE escrow row-copy routine" -- retained the whole-guest key and silently dropped the off-site
data key, which is the secret the retention was built to preserve. And because the copy happens
as the new blob overwrites the old, the destroying act was the ESCROW CEREMONY: the exact thing
a rebuilt box tells its customer to run, on a card promising in Hungarian that the old backups
stay recoverable. Both demo boxes crossed that line on 2026-08-04.
- identity_blob added to the table (CREATE + additive ALTER) and carried in the shared copy
routine, so BOTH callers are fixed at once: re-escrow and host-delete demotion.
- ListSupersededEscrow reads it back; store.HostEscrow gains IdentityBlob.
- CountCurrentEscrowWithIdentity is the census of who the fix protects.
- Nothing is backfillable: pre-v0.93.0 retained rows have no blob and their sources are gone.
- Tests assert the CONSEQUENCE (a retained row can still yield a repo password), which is why
the pre-existing retention test stayed green for two months asserting the mechanism.
R-197 — SaveHostEscrow returns the hash it replaced; the escrow PUT raises
offsite_repo_key_changed (warning, operator-only, edge-triggered) when both hashes are known and
differ. No hash value travels. Severity chosen for the world v0.93.0 creates: with the identity
blob retained, a changed key is "this history now depends on an older recovery code", not a loss.
R-192 (half) — the stuck alert now reports the two shapes it actually covers, burned and
regressed, each stating its own measurement; the regressed text withdraws the Re-issue
recommendation. Every self-heal refusal leaves a notification_log row with its reason. The
guard's logic is unchanged; its 500-oldest-reports scoping stays OPEN and the window is named in
the alert text so the limitation travels with the number. offsite_delivery_stuck and
offsite_credential_restaged are added to operatorOnlyEvents -- neither was registered and neither
has a customerMessages entry, which is not a block.
R-196 — five comments (not the three the spec expected) claimed ReissueCredentials rotates the
restic repo password. It resets the PROVIDER password and cannot touch the repo password, which
is generated on the box. All five corrected; the staleness mark documented as precautionary. The
BEHAVIOUR stays open.
Not in this release: R-199, R-200, R-201 remain open -- the chain that hands the key back is
still unassembled. Part 5 hit its gate; the orphan card is untouched (R-202).
Found by reading the first REAL digest, not by design. Every app row ended with
the same usage clause the mail already prints once on its own Filesystem line.
On a two-app box that is untidy; down a list of a dozen it is the same forty
characters twelve times, pushing the part that DIFFERS off a phone screen at
07:00 — the only moment this mail has to work.
The reserve's refusal message is authored for a single-app alert where naming
the filesystem is right, so the message is unchanged; the digest trims the
duplicate when rendering. trimRepeatedUsage removes ONLY an exact
"— <target path>:" suffix, so an unrelated reason is untouched and a reason that
is nothing but the usage clause is left alone rather than emptied.
Also updates TestRecoveryUnitCaptureFailed_NeverReachesTheCustomer, which
required the OPERATOR to be emailed a per-app capture failure. That was correct
when the event was the only signal and is wrong now that it is the record and
the digest is the notification. Its customer-safety claim is unchanged and is
why the test still exists; the operator assertion is inverted with the reasoning
written in place, and R-158's guarantee is shown to have MOVED, not weakened.
processOperator's cooldown no longer returns bare. It dropped the event BEFORE
LogNotification, so a suppressed operator alert and an event that never happened
were indistinguishable — from the operator's side and from the hub's own records.
Measured 2026-08-03: nine recovery_unit_capture_failed events arrived, two were
mailed, seven left no row anywhere. That is why the defect took a day to get the
right way round: there was nothing to read.
A suppressed operator event now writes a `suppressed` row carrying the message
and the key that suppressed it. This applies to EVERY operator event, not only
the one that exposed it. It does NOT change the cooldown's duration or semantics.
backup_run_failures: the per-run digest. In allowedEventTypes AND in
operatorOnlyEvents — allowlisting alone does not make an event operator-only,
and FormatCustomerEmail falls back to the raw English message rather than
blocking. A test demonstrates a customer with the type enabled receiving nothing.
recordOnlyEvents: a third routing class — stored and recorded, never mailed.
recovery_unit_capture_failed moves here: it is the record, the digest is the
notification. A register rather than downgrading severity to info, which would
relabel a genuine failure as informational everywhere it is queried.
cooldownRunSuffix: a sibling of cooldownTierSuffix, not a branch inside it, so
tier keeps byte-identical semantics and R-97a's tests are untouched. It makes
the cooldown effectively inert for the digest, which is the intent — a digest is
already rate-limited by construction; the refresh sweep sends no run_id and so
stays under the ordinary hourly cooldown.
The email renders as a list, not a JSON blob. An absent space reading renders as
unavailable, never as zeros.
New OPERATOR-ONLY event type recovery_unit_capture_failed (controller
v0.191.0, R-158): in allowedEventTypes AND notify.operatorOnlyEvents.
Deliberately not a reuse of backup_failed, which carries customer copy and
sits in the controller's DefaultEnabledEvents — reusing it would email the
customer in Hungarian about a failure they cannot act on. R-158's own
proposal said backup_failed; D-c overrides it.
disk_warning/disk_critical lose their generic customerMessages entries.
Both were allowlisted, copy'd, default-enabled and checkbox'd with NO
producer anywhere; controller v0.191.0 becomes that producer and sends a
DYNAMIC Hungarian message naming the drive and its free space.
FormatCustomerEmail prefers the entry over the message, so keeping a static
entry would discard the label and the byte figures — the same reason
offbox_enlarge_blocked and disk_health_degraded have none. The deletion is
pinned by a test.
New notify.IsOperatorOnly so the api package can pin BOTH registers of a new
event type in ONE test; allowlisted-but-not-operator-only is invisible when
they are checked separately, and it is the defect v0.78.0 shipped. The
register itself stays unexported.
REUSE.md's "new event type" extension point rewritten: it told readers to
always add a customerMessages entry, which is wrong for operator-only types
and harmful for dynamic-message ones.
Tests 574 -> 579. Red-proof: removing the operatorOnlyEvents entry shows the
customer being emailed; the skipped/operator_only row is asserted as a
positive observable.
An event type the hub does not allowlist makes POST /event return 400 and the
event vanishes (R-97a). The controller cannot emit backup_target_absent until
this is live, so the hub half ships first.
E-2 Phase 0 established that an absent backup target has NO prompt signal today.
The controller's drive-gate path stops apps and logs a WARN but emits nothing:
NotifyStorageDisconnected is defined and never called anywhere (verified against
the gitignored-cmd/ trap with a positive control). A drive that is ONLY a backup
target has no apps to stop, so it is entirely silent. The sole signal is the
tier's own failure at its next due cycle -- up to ~24h on the daily local tier,
which is the R-100 shape: a real fault visible only after a deadline elapses.
Added to BOTH registers, because each half fails differently:
allowedEventTypes -- without it the event is lost at the door;
customerMessages -- without it the event IS delivered but in the controller's
raw operator English, and nothing looks broken.
backup_target_absent is deliberately NOT folded into storage_disconnected: that
says "a drive went away and some apps may have stopped"; this says "the thing
that makes your backup survive a disk failure is gone".
Hungarian copy names the consequence, not just the fact. backup_target_restored
is the paired recovery at info severity -- severityNotifies NOT widened.
Three tests pin the pair and the copy's substance. All red-proofed with the
mutation VERIFIED to have landed first: the initial attempt silently no-op'd
(gofmt had realigned the map) and the test "passed" -- a false proof that would
have been reported as evidence.
Green gate: build + vet + test rc=0, run separately from this commit.
v0.78.0 asserted in a comment that a type with no customerMessages entry cannot
reach a customer. It can: templates.go falls back to the raw message when the
entry is missing, and the only customer gate is prefs.EnabledEvents — pure
configuration. A customer with whole_guest_backup_failed enabled would have been
emailed raw English operator text about a backup they cannot act on. The new test
proves it against the v0.78.0 shape.
operatorOnlyEvents is now an explicit register checked before prefs, logging a
skipped/operator_only row so the skip is visible. NOT implemented as 'missing
customerMessages blocks delivery' — several types rely on that fallback on
purpose. The handler comment now names the real mechanism.
internal/quiesce had no route to the hub at all: three failed whole-guest backups
on 2026-07-27 produced zero events. Hub half of the fix.
whole_guest_backup_failed / _recovered are allowlisted with NO customerMessages
entry. Deliberately not backup_failed/backup_completed — those have customer
Hungarian templates AND sit in demo-felhom's live enabled_events, so reusing them
would email the customer that their backup failed while it is still retrying
behind the R-88 breaker.
The recovery joins recoveredPairedDownTypes because it is severity info and
severityNotifies drops info — otherwise the operator hears it break and never
hears it heal. Its customer leg is pairing-gated and can never fire.
Operator cooldown gains a per-tier dimension from the event details, so one tier
cannot mask another for an hour. Narrow: empty suffix unless a tier is sent, so
no existing event type changes.
Let a customer bind their own freshly-installed appliance without the
operator: operator "Send self-bind link" mints a 7-day tokenized
capability link, emailed (Hungarian, sibling sender) to the customer, who
opens a public /bind/<token> page and proves two factors — the console
pairing code shown on the box screen + their retrieval passphrase — and
the hub stages the bind via the same BindAppliance (provenance
customer_selfbind). The box's ~30s appliance poll delivers.
Viktor's three rulings verbatim: console pairing code (no appliance list
ever rendered), operator-sent tokenized link, 5-attempt lockout ->
"call support". Wrong code == wrong passphrase (one generic failure, no
oracle, both factors compared unconditionally); expiry falls back to
operator-bind unchanged.
THE TRAP: one public prefix /bind/, exempt from auth+CSRF at both /login
gate sites via a single isPublicBindPath predicate (tight trailing-slash
match; ServeMux ..-cleans; handler rejects '/' in token). 9 tests
(Scenarios A-F + F1/F2); 4 red-proofs verified red-then-green (lockout,
oracle, widened-prefix, single-active). GC verdict: no appliance GC ->
the 7-day TTL stands alone. Controller/agent untouched; R-27b deferred.
Green: full hub build/vet/test (17 ok) + bash -n + hub confirm gate.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017qDiBqKKQ5vPB5fXBqu7Kp
Makes PBS DR storage visible like the restic pool box (v0.64.0), differentiated. Scoping
correction: restic = subaccounts on the shared Hetzner Storage Box (Hetzner API); PBS DR =
the felhom-offsite PBS datastore on the ep0 endpoint VM (NO Hetzner API). Option A
(Viktor-ruled): a read-only `usage` op on the felhom-tenantsync ep0 forced command (twin of
fingerprint), polled by a new hub checker on the 15-min throttle. READ-ONLY throughout.
Phase-0 (gate PASSED): on ep0 (PBS 4.2.3), df -B1 --output=size,used,avail <datastore path>
yields bytes (39990112256/7627939840/... ~19%), read-only, existing sudo context, no admin token.
- scripts/felhom-tenantsync.sh -> v1.2.0: read-only `usage` short-circuit (df on the datastore
path), no customer_id, no admin token, NO mutation. + a bash harness proving zero mutation.
- tenantsync.Client.Usage() + BoxUsage; unknown-op -> typed ErrUsageUnsupported (graceful).
- monitor.PBSDRBoxChecker: OffsiteBoxChecker clone over a usageReader seam; 15-min throttle,
cached PBSBoxSnapshot, escalation-only pbsdr_box_fill on the "pbsdr-box" scope (operator only,
no SaveEvent), recovery re-arm. Fill only. THREE states: ok / unavailable (ep0 <=v1.1.0,
neutral no-alert) / degraded (exec failed, keep last).
- config: Alerting.PBSDRBoxFill{Warn,Crit}Percent (80/90); built with the tenantsync client,
60s sweep, SetPBSDRBox. Hub deploy INDEPENDENT of the ep0 update (graceful degradation).
- web: /offsite splits into Restic + PBS DR hash tabs (endpoint cards under PBS DR); PBS panel;
the single dashboard tile becomes two gauges (RESTIC pct.ratio, PBS DR pct / n/a).
- runbook offsite-endpoint.md 10: v1.2.0 update steps (no sudoers/authorized_keys change).
Tests: 10 Go + the harness; 3 red-proofs (usage mutation, escalation-only, unavailable-drives-band)
confirmed red then restored. go build/vet/test + bash -n + hub confirm gate all pass.
The operator sees the shared pool box's real state on the hub: total box fill vs
capacity, Σ(shared soft quotas) vs capacity (the oversubscription ratio), per-customer
usage/quota bars, and a box-level operator alert (fill % + oversub ratio) on the existing
dispatcher's operator channel. Per-customer fill alerts already existed; the box-level
aggregate was the gap. READ-ONLY against Hetzner (GET only).
Phase-0 probe (gate PASSED): the live pool box 611714 returns capacity via
storage_box_type.size (1 TiB / bx11) and usage via a stats object (size/size_data/
size_snapshots), all bytes; our token reads it (200).
- hetznerapi: additive StorageBoxType + StorageBoxStats on StorageBox (no existing field/
method changed); fake carries them + a GetBoxCalls counter; golden decode test.
- monitor.OffsiteBoxChecker: OffsiteChecker-sibling for the box; fetch-throttled (1 GET/
15min), cached BoxSnapshot, escalation-only + recovery re-arm. FILL (used/capacity 80/90)
+ OVERSUB (Σ shared+enabled quotas / capacity, 2.0x) — independent. Σ from the ConfigJSON
Descriptor (offsite.ReadDescriptor, new), never the report echo; dedicated+disabled
excluded. Scope "pool-box" -> operator channel only, no SaveEvent. Failed fetch keeps the
last snapshot degraded; missing data never becomes 0% and never transitions a band.
- config: Alerting.OffsiteBoxFill{Warn,Crit}Percent + OffsiteOversubWarnRatio (80/90/2.0
defaults; thresholds pending Viktor's ruling). Constructed in the HETZNER_TOKEN branch,
60s sweep, snapshot handed to the web server.
- web: Offsite-tab panel (fill bar, Σ+ratio, per-customer usage/quota rows) + a compact
dashboard tile; reads the cached snapshot only, never fetches; nil -> "not configured".
Tests: 10 new + 4 red-proofs (throttle, Σ filter, escalation-only, failed-fetch honesty),
all confirmed red then restored. go build/vet/test all pass; hub confirm gate OK.
ProcessEvent routed only warning/error; a critical-severity event was silently dropped.
Now routes warning/error/critical, logs unrecognized severities, and guards a nil
GetNotificationPrefs (which would panic/crash the hub). host_disk_critical emits its
natural critical severity.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HxLA1mZurFq9kt8hneFeCs
New HostDiskChecker on the 60s sweep alerts the operator when a Proxmox host root
filesystem crosses warn (90%) / crit (95%). Born/persistent (a disk already full at
hub restart alerts on cycle 1); distinct host_disk_* event types from the guest disk_*;
critical band maps to severity error (the dispatcher only routes warning/error).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HxLA1mZurFq9kt8hneFeCs
Restores notify/templates.go, store/telemetry.go, web/configs.go to upstream —
those were alignment-only churn from a tree-wide gofmt, not part of slice 3. Keeps
the host-domain diff additions-only.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>