Commit Graph

181 Commits

Author SHA1 Message Date
admin 63f29c6ad8 hub v0.113.0: deploy (R-497 passphrase hand-over copy)
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-14 17:28:53 +02:00
admin 2f5d3af6f9 deploy: hub 0.112.0 (R-472 declared MinAgent floor)
gates / gates (push) Successful in 20s
2026-09-13 16:49:15 +02:00
admin 0f65f7a197 manifests: hub 0.111.0 -> 0.111.1 (R-434, the alarm's withdrawn promise)
gates / gates (push) Successful in 18s
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-09-01 18:35:58 +02:00
admin 65c82c4aa0 manifests: hub 0.110.0 -> 0.111.0 (R-431)
gates / gates (push) Successful in 16s
The manifest is the truth - a code push and an image build deploy NOTHING until this tag changes in
git AND the app is synced. Auto-sync is OFF.
2026-09-01 14:26:38 +02:00
admin 1aeaa30c28 hub v0.110.0: allowlist + operator-only for offsite_proof_empty (R-87)
gates / gates (push) Successful in 15s
Controller v0.231.0 adds a nightly job that proves an app's newest off-site snapshot still
CONTAINS that app's data. When it finds one that does not, it emits offsite_proof_empty.

Two register lines, both load-bearing and both in this commit:
- allowedEventTypes - an unallowlisted type is answered 400 and VANISHES, so this entry is
  what makes the alarm exist at all.
- operatorOnlyEvents - a missing customerMessages entry is NOT a routing block
  (FormatCustomerEmail falls back to the raw message, the v0.78.0 defect that register was
  built for). A customer can take no action on a hollow recovery unit.

DELIBERATELY NOT reusing backup_integrity_failed, which is the nearest existing type: it
means THE STORE IS DAMAGED and carries the Hungarian template saying so. Here the store is
sound and the CONTENT is absent - different cause, different action, and telling a customer
their backups are damaged when they are not is the more expensive mistake. Same asymmetry
looksLikeRepositoryDamage is shaped around.

DELIBERATELY no customerMessages entry (the controller's dynamic Hungarian names the app and
what is missing; a template would discard it) and DELIBERATELY not in perAppCooldownEvents
(a fenced act - this job proves ONE app per night, so the coarse hourly cooldown is already
the right grain).

This widens the R-87 task's stated repo scope to felhom.eu/hub/. The reason is recorded in
felhom-controller/CONTEXT.md ruling 4 rather than left as an unexplained diff.

Hub green gate: go build/vet/test all pass, 18 packages.
2026-08-31 20:53:13 +02:00
admin f5c9411e5e R-331 (hub half): the Backup card reads offsite, not the dead backup fields (v0.109.0)
The customer page's Backup card read `Snapshots 0 / Repo Size 0 MB / Integrity
Unknown` for EVERY customer, indefinitely. Measured on demo-hp 2026-08-30 while
that night's controller log said `[offbox] backup OK: 8 app(s) backed up, 67
snapshot(s), 2m14s` and the box held snapshot_count:67, repo_size_bytes:
140829678, stats_known:true.

A card reading "no backups" over a working backup is worse than no card -- the
R-88 direction of failure (degrade to NO BACKUP rather than to UNKNOWN) on the
one screen that answers "is this customer protected?".

The data was never missing. The card rendered the report's `backup` object,
whose snapshot/size/integrity fields have had no producer since slice 8C. The
live numbers are in the `offsite` object, which THIS PACKAGE already reads for
the Offsite page and which monitor.OffsiteChecker already alarms from. Proof the
bytes were arriving: the Offsite page rendered demo-hp's usage as 0.1 GB from
that very object while the Backup card said 0 MB. So this is a render fix over
an existing feed, not a new pipeline.

Not a one-line swap, because snapshot_count:0 means two opposite things --
"holds nothing" and "never measured". R-225 measured that confusion one layer
down. backup_card.go resolves a three-way ruling in Go (a {{if}} chain over
map[string]interface{} float64s cannot keep the absent/zero distinction the card
is entirely about):
  no offsite object   -> "No off-site data reported", and says explicitly that
                         this is NOT the same as "no backups"
  disabled + state    -> names the blocker (needs_credential)
  stats_known:false   -> em-dash + "never been measured". NEVER 0
  stats_known:true    -> the real numbers, INCLUDING a real 0

A pre-v0.225.0 controller sends no stats_known -> false -> "unknown". That
direction is pinned: upgrading the hub ahead of the fleet must not report every
un-upgraded customer as having zero backups.

The Integrity row is DELETED, not re-sourced: nothing produces it, the
controller runs no integrity check, and NotifyIntegrityOK/Failed are called from
nowhere.

RED-PROOF: restore the pre-fix card markup -> all four tests fail, reporting 67
and 134.3 MB absent from the rendered page and the Integrity row present. The
tests drive handleCustomerUnified and grep the HTML on purpose: the defect was
the template's choice of source object, so a test one layer below it would have
been green against the shipped bug.

Green gate clean: 18 packages, rc 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-08-30 18:39:15 +02:00
admin 45659bdc5a hub v0.108.0: deploy the per-app cooldown grain (R-389)
gates / gates (push) Successful in 17s
Image built and pushed to the registry BEFORE this manifest bump lands, so a
sync can never point at a missing tag. Auto-sync is off; the sync that follows
is deliberate. Never kubectl set image.
2026-08-23 13:54:48 +02:00
admin 68a9f5475c hub v0.107.0: the hub rewrote a severity and said nothing (R-387); golden 0.223.0
gates / gates (push) Successful in 17s
One handler, two fields, opposite discipline. An unknown event_type is rejected
with a loud 400. An unknown severity was rewritten to "info" without a word -
and severityNotifies drops "info" before BOTH legs, so the event was stored,
answered 200, and mailed to nobody.

Two shipped features went out that way: DiskAlertKind.Severity emitted "warn"
until controller v0.215.0, app_start_failed until v0.223.0. Measured on the live
hub DB today: 91 app_start_failed events stored all-time, ZERO notification_log
rows before this session - not one, on any channel.

The mechanism built to catch this class was structurally blind to it: the
dispatcher's `unrecognized severity` line cannot execute for anything arriving
over the API, because the coercion one line earlier guarantees the value it
looks for cannot arrive.

The coercion STAYS - a rejected event is a lost event, and losing an alarm is
worse than mis-routing one. Only the silence is fixed: a WARN naming the
customer, the event type and the rejected value.

The dispatcher branch is KEPT, not deleted as dead, and the reason is evidence
rather than caution: cmd/hub/main.go wires dispatcher.ProcessEvent DIRECTLY as
the monitor.EventNotifyFunc for the staleness, host-staleness and offsite-box
checkers, which never pass through the handler. For those it is the only
severity guard there is. All 90 severity literals in internal/monitor are
already valid, so the guard is silent because the producers are correct.

Test count 702 -> 709. Red-proof seen failing: delete the WARN line and the
coercion test fails with "the hub rewrote a severity and said nothing".

Golden 0.223.0 baked and published (sha 9eaf39ac3921...), round-trip HTTP 206.
Vouching is the operator's act and was not done here.
2026-08-23 11:57:26 +02:00
admin c03f629d43 manifests: hub 0.105.0 -> 0.106.0 (R-339 box reachability)
gates / gates (push) Successful in 14s
The image is built and pushed; this is the change that actually deploys it.
A code change plus a CHANGELOG bump deploys nothing -- the running image
moves only when this tag moves in git and the app is synced.
2026-08-18 19:28:46 +02:00
admin bbd59f4a44 Deploy hub 0.105.0 — the third name, the quiet-machine alarm, and the copy guard
gates / gates (push) Successful in 13s
2026-08-13 15:51:49 +02:00
admin 7c97c949f6 Deploy hub 0.104.0 — the guest-network reader and the naming half
gates / gates (push) Successful in 25s
2026-08-13 10:51:23 +02:00
admin 823cd2949b Publish installer-v1.28.0: move both git-sync refs to the new tag
gates / gates (push) Successful in 13s
Pushing publishes nothing here - /scripts/ follows the installer TAG, and both the
sidecar and the init container carry the ref. Verified against the live URL after
the sync, because the runbook still says otherwise (R-309, still open).
2026-08-13 08:16:28 +02:00
admin 8b188bea68 hub v0.103.0 — a host can read the packages we kept for it (R-311)
gates / gates (push) Successful in 37s
ListSupersededEscrow had zero production callers for nineteen days. It is the only
reader of a retained identity_blob, so the retention shipped in v0.93.0 was
material the product could not reach - proven on the fixture 2026-08-12, where a
code that opens a retained package was answered as a code that opened nothing.

New GET /api/v1/hosts/<id>/escrow/retained: self-scoped exactly as the current-row
GET, same recovery-mode gate, same audit event written BEFORE the bytes leave,
capped at 16. Rows with a NULL identity_blob are WITHHELD and returned as
unopenable_count - they retain the PBS key, not the repository password, so they
can never open what the caller is asking about, and serving them would let the
screen claim an earlier package is openable on exactly the boxes the original
defect hurt. The count is returned because their existence is load-bearing and
underivable by the caller.

The trade, stated rather than waved through: the hub still cannot read any of it -
sealed bytes in, sealed bytes out, no decrypt path, no recovery code ever held.
What widens is volume, bounded by self-scope, the recovery-mode gate and the cap.

The response is a NAMED TYPE, not a map, so the wire-contract gate can resolve it;
the wire is declared as a fourth ROOT and the gate now checks 182 tags rather than
174. A positive control shows that check is name-presence, not decodability -
filed as R-315 rather than reported as coverage.

Also: golden 0.214.0 baked, published and round-trip verified; the countdown on
demo-felhom cancelled on the operator's ruling (R-307); the spike that halted
Part 3 recorded as R-312; the set-aside store found unrecoverable as R-313.
Six hub tests through the real endpoint; four red-proofs asserted applied.
2026-08-12 18:56:44 +02:00
admin 1d5f2b8bb6 DRILL: the retained key works, and the customer cannot reach it
gates / gates (push) Successful in 23s
Three verdicts, kept separate because collapsing them is how this assumption
survived a week.

(a) The material IS retained. host_escrow_superseded id 11 is the first retained
row in fleet history to carry identity_blob (572 B), byte-identical to the
pre-supersession row (sha256 a10032341c8584ed...).

(b) The retained material DOES open the old store. Unsealed with the old recovery
code it yielded a password byte-identical to the pre-change one, and restored
three planted files byte-identical from a store the box itself could no longer
open - including a Hungarian accented filename verified as raw bytes. Negative
control ran first and failed closed.

(c) The customer has NO route, and is misinformed. ListSupersededEscrow has zero
production callers; the recovery path selects FROM host_escrow. Asked with the
code that had just worked by hand, the product answered "the recovery code did
not open the sealed bundle". A valid code for retained history is reported as a
bad code - the R-224 class again. R-304, rank 1.

Both installer faults were watched happening first, so installer-v1.27.0 is now
published (tag + both webpage.yaml refs). Pre-fix: the box came up on controller
0.98.3 against a vouched 0.213.0, below the floor and below the version carrying
the recovery screen; and our own uninstall left dnsmasq on 0.0.0.0:53 so our own
next install refused. R-297 and R-300 CLOSED.

Also filed R-305 (the dnsmasq fix fires once per machine - the leftover returns
on the second reinstall, proven), R-306 (--preflight-only writes state it says it
does not), R-307 (a live abandon countdown on demo-felhom, firing 2026-08-24 -
operator decision), R-308 (stored controller password stale), R-309 (the day-0
runbook's publication claim has been false since R-110), R-310 (two edges).

Ceiling R-303 -> R-310. Capability map moved: the retention claim is now marked
operator-only. Phase A logs did not survive the intermediate revert; recorded.
2026-08-12 17:41:56 +02:00
admin 36bcd12543 Deploy hub v0.102.0, and record that the package deleter is still not established
gates / gates (push) Successful in 21s
manifests/hub.yaml 0.101.0 -> 0.102.0. Image built and pushed, and verified
served by the registry before the bump rather than after.

R-287 corrected on two counts. My own sentence "no DELETE on the packages API
appears in 48h of Gitea router logs" is WITHDRAWN: kubectl logs on the Gitea pod
now returns nothing older than 2026-08-09 16:35 and contains zero api/packages
lines even for requests I made myself, so the log never covered the window and
its silence was never evidence.

A second attempt to attribute the deletion also failed and the deleter remains
NOT ESTABLISHED. Sources exhausted: no register row records a package prune
(R-210 is WAITING-ON-OPERATOR, says "Nothing was deleted; this is a list, not an
action", and concerns local Docker images); package_version has no soft-delete
column so a deletion leaves no row; Gitea's action feed carries no package
operation at all in the window; and a uniform newest-ten cap is not visible --
felhom-agent generic holds 10 but the container packages hold 19 each. It may be
unestablishable from this side: Gitea keeps no package-deletion trail.
2026-08-09 19:15:17 +02:00
admin d7d147ed5a deploy: hub 0.101.0 (60s artifact-dropdown memo)
gates / gates (push) Successful in 29s
2026-08-08 20:10:03 +02:00
admin 047e296ad1 deploy: hub 0.100.2 (connection reuse for the Gitea client)
gates / gates (push) Successful in 31s
2026-08-08 17:56:34 +02:00
admin de0110b8da deploy: hub 0.100.1 (the side-by-side dropdown resolve rides this image)
gates / gates (push) Failing after 14m52s
2026-08-08 17:51:29 +02:00
admin 5ef7b92b67 deploy: hub 0.100.0 — the manifest is the truth
gates / gates (push) Successful in 25s
felhom-hub:0.100.0 confirmed present in the registry (manifest HTTP 200) before this bump. A built
image deploys nothing until this tag moves in git and the app is synced.
2026-08-08 17:43:07 +02:00
admin 9771fd9c27 deploy: hub 0.99.0 — the manifest is the truth (R-260)
gates / gates (push) Failing after 18s
A built image deploys nothing until this tag moves in git and the app is synced.
felhom-hub:0.99.0 confirmed present in the registry (manifest HTTP 200) before the bump.
2026-08-08 09:03:23 +02:00
admin b12f8ec2f3 deploy hub v0.98.0 — the R-241 superseded-package purge
gates / gates (push) Successful in 11s
The manifest is the truth: the image was built and pushed, but nothing
deploys until this tag moves and the app is synced.
2026-08-07 12:20:59 +02:00
admin 79e31ac24b manifests: hub 0.97.0 -> 0.97.1 (held-floor reason matches its cause)
gates / gates (push) Successful in 9s
2026-08-05 17:54:33 +02:00
admin cad0406e2b manifests: hub 0.96.0 -> 0.97.0 (R-216 floor guard, R-222 ACK fields)
gates / gates (push) Successful in 7s
2026-08-05 17:50:23 +02:00
admin 4114c5f891 manifests: hub 0.96.0 (R-204 item 4, hub half)
gates / gates (push) Successful in 7s
2026-08-05 10:51:44 +02:00
admin 975a690fbe manifests: hub 0.95.0 (R-196 / R-204 item 2)
gates / gates (push) Successful in 7s
2026-08-05 07:19:57 +02:00
admin dd089265e8 manifests: hub 0.93.0 -> 0.94.0 (R-199 box-authenticated escrow retrieval)
gates / gates (push) Successful in 8s
2026-08-04 13:40:26 +02:00
admin 40687b0921 manifests: hub 0.92.0 -> 0.93.0 (R-198 escrow retention + R-197/R-192 honesty pass)
gates / gates (push) Successful in 7s
2026-08-04 12:58:04 +02:00
admin f581ac1349 manifests: hub 0.91.1 -> 0.92.0 (R-195 phantom-customer alarm)
gates / gates (push) Successful in 6s
2026-08-04 11:05:41 +02:00
admin 71662336aa manifests: /scripts/ syncs installer-v1.24.0 -> installer-v1.25.0 (R-191)
gates / gates (push) Successful in 7s
2026-08-04 09:46:21 +02:00
admin 311dc06c13 manifests: /scripts/ syncs installer-v1.23.0 -> installer-v1.24.0 (R-185)
gates / gates (push) Successful in 8s
Both refs — the git-sync sidecar and its init container. A fresh pod must not
serve a different installer from a running one.
2026-08-03 18:59:05 +02:00
admin ff2655cf19 manifests: hub 0.91.0 -> 0.91.1 (R-86: observation may only widen a tier's window)
gates / gates (push) Successful in 7s
2026-08-03 15:18:00 +02:00
admin 687fedd8ee manifests: hub 0.90.1 -> 0.91.0 (R-86 Part 2, per-tier restore-proven window)
gates / gates (push) Successful in 7s
2026-08-03 15:07:52 +02:00
admin f21e7caed1 hub v0.90.1 — the digest's per-app lines stop repeating the filesystem figures (R-182)
gates / gates (push) Successful in 7s
Found by reading the first REAL digest, not by design. Every app row ended with
the same usage clause the mail already prints once on its own Filesystem line.
On a two-app box that is untidy; down a list of a dozen it is the same forty
characters twelve times, pushing the part that DIFFERS off a phone screen at
07:00 — the only moment this mail has to work.

The reserve's refusal message is authored for a single-app alert where naming
the filesystem is right, so the message is unchanged; the digest trims the
duplicate when rendering. trimRepeatedUsage removes ONLY an exact
"— <target path>:" suffix, so an unrelated reason is untouched and a reason that
is nothing but the usage clause is left alone rather than emptied.

Also updates TestRecoveryUnitCaptureFailed_NeverReachesTheCustomer, which
required the OPERATOR to be emailed a per-app capture failure. That was correct
when the event was the only signal and is wrong now that it is the record and
the digest is the notification. Its customer-safety claim is unchanged and is
why the test still exists; the operator assertion is inverted with the reasoning
written in place, and R-158's guarantee is shown to have MOVED, not weakened.
2026-08-03 13:54:02 +02:00
admin dd40f85bb8 hub v0.90.0 — a dropped notification leaves a trace, and the backup digest arrives (R-182)
gates / gates (push) Successful in 7s
processOperator's cooldown no longer returns bare. It dropped the event BEFORE
LogNotification, so a suppressed operator alert and an event that never happened
were indistinguishable — from the operator's side and from the hub's own records.
Measured 2026-08-03: nine recovery_unit_capture_failed events arrived, two were
mailed, seven left no row anywhere. That is why the defect took a day to get the
right way round: there was nothing to read.

A suppressed operator event now writes a `suppressed` row carrying the message
and the key that suppressed it. This applies to EVERY operator event, not only
the one that exposed it. It does NOT change the cooldown's duration or semantics.

backup_run_failures: the per-run digest. In allowedEventTypes AND in
operatorOnlyEvents — allowlisting alone does not make an event operator-only,
and FormatCustomerEmail falls back to the raw English message rather than
blocking. A test demonstrates a customer with the type enabled receiving nothing.

recordOnlyEvents: a third routing class — stored and recorded, never mailed.
recovery_unit_capture_failed moves here: it is the record, the digest is the
notification. A register rather than downgrading severity to info, which would
relabel a genuine failure as informational everywhere it is queried.

cooldownRunSuffix: a sibling of cooldownTierSuffix, not a branch inside it, so
tier keeps byte-identical semantics and R-97a's tests are untouched. It makes
the cooldown effectively inert for the digest, which is the intent — a digest is
already rate-limited by construction; the refresh sweep sends no run_id and so
stays under the ordinary hourly cooldown.

The email renders as a list, not a JSON blob. An absent space reading renders as
unavailable, never as zeros.
2026-08-03 13:46:48 +02:00
admin bee6848458 installer v1.23.0 — publishing becomes an act, not a side-effect (R-110, R-183)
gates / gates (push) Successful in 8s
Two channels moved off main in the same change, because either one left behind
makes the other cosmetic.

Channel 1 — the served script. webpage.yaml git-synced /scripts/ from
--branch=main every 30s and nginx served that tree, so pushing this file WAS
publishing it: within half a minute it was what every new machine downloaded and
ran as root, with no staging and no rollback but another push. The sync is now
SPLIT: the website keeps tracking main at the same cadence (a copy edit must
never need a release) and /scripts/ tracks the tag installer-v<SCRIPT_VERSION>.

PROVEN before the manifest was touched: git-sync v4.4.0 follows a tag AND
notices a MOVED one — measured on a throwaway sync against this repo,
"update required ... local:<old> remote:<new>" -> "updated successfully",
within one period. The moved-tag half is what the publish model rests on.

Channel 2 — the sixteen files fetched at run time. fetch_raw pulled from
$AGENT_REPO/raw/branch/main; it now pulls raw/tag/v$ART_AGENT_VER. That is a
correctness fix, not only a channel one (R-183): a fresh install fetched the
vouched agent BINARY while taking its unit file, sudoers and guarded wrappers
from whatever main held. Two refs, one install, nothing compared them. Their
correct ref was never SCRIPT_VERSION — they do not live in this repo.

No fallback to a branch: a vouched version whose tag is missing fails loudly
rather than quietly serving main.

Channel 3 — the URL — needed no change, recorded rather than left silent:
https://felhom.eu/scripts/felhom-host-install.sh never carried a ref, so both
producers follow the tag with no edit. No hub change, no hub version bump.

Gate 6 in hostinstall_gates.py pins all three structurally with no network, so
it stays in --fast and runs in CI. It deliberately does NOT assert "a tag exists
for the current SCRIPT_VERSION": that would go red on the very push that bumps
the version, before publishing — and publishing being separate is the ruling.
2026-08-03 12:08:37 +02:00
admin 6d359a5360 manifests: hub 0.88.0 -> 0.89.0 (R-167/R-158 event routing)
gates / gates (push) Successful in 7s
2026-08-02 23:21:17 +02:00
admin d5774d3189 manifests: hub 0.87.0 -> 0.88.0 (R-172 WAL fix)
gates / gates (push) Successful in 8s
2026-08-02 21:08:17 +02:00
admin 8d9b78c153 manifests: hub 0.86.0 -> 0.87.0 (R-94, the Setup tab renders no version) 2026-08-02 15:29:51 +02:00
admin 80f473999e manifests: hub 0.85.0 -> 0.86.0 (Copy without reveal) 2026-07-31 09:23:12 +02:00
admin 37f7ff69f2 manifests: hub 0.84.0 -> 0.85.0 (Network card) 2026-07-31 08:50:42 +02:00
admin edc7dcc9e2 manifests: hub 0.83.0 -> 0.84.0 (Console access card) 2026-07-31 08:20:48 +02:00
admin acfc2b7e95 R-109 + R-122: the recipe assembly stops dropping sections (hub v0.83.0)
AssembleDRRecipe's hostHalfShape/appHalfShape are ALLOW-LISTS, not the
forward-compat their comment advertised: a section an emitter adds is silently
discarded until it is named in both the shape struct and AssembledRecipe. No
error, no log, no failing test.

R-122 (found this session): that already happened and shipped. The controller
has emitted offsite_restic since fork-4 — the offsite recovery LOCATION — the
hub stored it for all three real customers, and appHalfShape never listed the
key, so no delivered recipe has ever contained it. It stayed green because the
fixture drAppHalf is hand-written and omits the field.

R-109: the agent's new backup_target is a new top-level host-half section and
would have been dropped identically, making the fix read as shipped while
changing nothing an operator can see.

3 tests built on halves read verbatim out of the live dr_recipe table, plus
2 red-proofs (each mutation asserted to have landed). vet rc=0, suite rc=0, 17 ok.

Registers: R-106 + R-109 dispositioned; R-105/R-106 were READY in ROADMAP with
no OPEN-ITEMS row (→ R-123, registered); R-124 filed on the "root" spelling.
2026-07-30 13:13:56 +02:00
admin 1a68b53b06 hub v0.82.0 (R-120): the vouch path REFUSES a golden the fleet has already outrun
The golden's version IS the controller it bakes (build-golden.sh:345 defaults
GOLDEN_VERSION to the controller tag), so a golden behind the newest deployed
controller means every FRESH install lands on stale application code. On the R-120
occurrence that stale code shipped a customer-facing falsehood: a box from the
0.185.1 golden told a customer whose backup drive had fallen out that the backup was
on the same disk as the system -- false, the drive was gone -- and offered a
different drive as the remedy.

WHY A GATE, NOT A REMINDER. The gap has opened three times: R-111 (golden's agent 17
releases behind), R-115 (agent built and deployed, never published), R-120 (this).
The first two were closed by re-baking and remembering; remembering then failed
again. R-29 is the standing proof that a check nobody runs is worse than none because
it reads as coverage -- hostinstall_gates.py sat RED and uninvoked across three
version bumps and hub_confirm_gate.py has never run at all. So the property that
matters is not whether a check exists but whether it BLOCKS.

- Wired into handleSetArtifacts (internal/web/configs.go), immediately before the
  only write, on the sole UI path to SetArtifactManifest -- it runs on every vouch
  without anyone choosing to. A script in scripts/ would have been a fourth orphan.
- It REFUSES (operator ruling, 2026-07-30), with a flash naming the remedy.
- Signal: store.NewestReportedControllerVersion() over reports.controller_version,
  SEMVER-compared in Go -- MAX() in SQL ranks 0.99.0 above 0.186.0, a pair this
  fleet has shipped. No outbound call, no new credential.
- Fail-open in exactly two deliberate cases: an empty golden field (clearing the
  manifest is legitimate) and an unknown fleet version (a new hub must vouch its
  first golden).

NEAR-MISS RECORDED: the first draft read guests.controller_version, a column that
exists in the schema and that NOTHING writes -- it would always have seen "" and
failed open, i.e. inert, this gate's own failure shape. Caught by grepping for a
writer before trusting the column.

Blind spot stated rather than papered over: a controller no box has ever run is
invisible to this signal. Not the failure that has bitten -- all three instances were
deployed-newer-than-baked.

4 tests through the PRODUCTION handler over httptest, never an injected seam. The
refusal asserts both the flash and that the manifest was NOT written, because a gate
that redirects and saves anyway reads as enforcement while providing none. Red-proof:
deleting the block makes the stale golden vouchable and both assertions fail.

ROADMAP R-29's audit list now records this as the FIRST enforced gate, so the
contrast with its three orphans is kept rather than lost. The orphans are unchanged.

Suite rc=0 read separately from this commit.
2026-07-30 10:42:43 +02:00
admin 6869a14015 manifests: hub 0.80.0 -> 0.81.0 (E-2 backup_target_absent event types)
The manifest tag is what ArgoCD deploys; the code commit and CHANGELOG bump
deploy nothing on their own. Ships BEFORE the controller: an event type the hub
does not allowlist is answered 400 and the event vanishes.
2026-07-29 08:27:44 +02:00
admin 4f34a9e0ae manifests: hub 0.79.0 -> 0.80.0 (R-100) 2026-07-28 13:19:41 +02:00
admin 2c0e43e0d0 hub v0.79.0 — R-97c: make the operator-only claim true
v0.78.0 asserted in a comment that a type with no customerMessages entry cannot
reach a customer. It can: templates.go falls back to the raw message when the
entry is missing, and the only customer gate is prefs.EnabledEvents — pure
configuration. A customer with whole_guest_backup_failed enabled would have been
emailed raw English operator text about a backup they cannot act on. The new test
proves it against the v0.78.0 shape.

operatorOnlyEvents is now an explicit register checked before prefs, logging a
skipped/operator_only row so the skip is visible. NOT implemented as 'missing
customerMessages blocks delivery' — several types rely on that fallback on
purpose. The handler comment now names the real mechanism.
2026-07-27 17:54:47 +02:00
admin 331193b898 hub v0.78.0 — R-97a: whole-guest backup events, operator-only
internal/quiesce had no route to the hub at all: three failed whole-guest backups
on 2026-07-27 produced zero events. Hub half of the fix.

whole_guest_backup_failed / _recovered are allowlisted with NO customerMessages
entry. Deliberately not backup_failed/backup_completed — those have customer
Hungarian templates AND sit in demo-felhom's live enabled_events, so reusing them
would email the customer that their backup failed while it is still retrying
behind the R-88 breaker.

The recovery joins recoveredPairedDownTypes because it is severity info and
severityNotifies drops info — otherwise the operator hears it break and never
hears it heal. Its customer leg is pairing-gated and can never fire.

Operator cooldown gains a per-tier dimension from the event details, so one tier
cannot mask another for an hour. Narrow: empty suffix unless a tier is sent, so
no existing event type changes.
2026-07-27 16:59:05 +02:00
Claude Code 73889e9fdf manifests: pin hub 0.77.0 (R-85 restore-test signals) 2026-07-27 07:34:43 +02:00
Claude Code 48daa4fdeb manifests: pin hub 0.76.0 (R-82 Slice C tier-aware thresholds) 2026-07-26 16:59:38 +02:00
Claude Code 88b41ec870 manifests: pin hub 0.75.0 (R-81 anchored backup deadline check) 2026-07-26 11:45:14 +02:00