Commit Graph

200 Commits

Author SHA1 Message Date
admin 99c0709cbe hub v0.123.0: app_stopped_unhealthy (decision 28); 09 decisions 26-28 recorded
gates / gates (push) Successful in 31s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 11:38:13 +02:00
admin e4d45a8f72 hub v0.122.0: app_hold_no_whole_copy — allow-listed, operator-only, per-app cooldown (R-659)
gates / gates (push) Successful in 28s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 07:33:04 +02:00
admin 511969928d hub v0.121.0: app_oom_storm allow-listed, operator-only, per-app cooldown (R-636)
gates / gates (push) Successful in 24s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 17:17:12 +02:00
admin d3b284863e hub v0.120.0: app_update_undone/held reach the household, per app, in its language
gates / gates (push) Successful in 26s
Allowlisted, not operator-only, seeded for new households and added
once (add-only) to every existing enabled_events row. mail.event entries
in hu and en name the app from details.stack_name. Per-app cooldown on
both the operator and the household leg.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-23 13:47:13 +02:00
admin e02bc03819 hub v0.119.0 — English households get English words for their codes (R-597); R-596/R-598 closed
gates / gates (push) Successful in 24s
The setup code and the owner passphrase now follow the household's language,
one word longer in English so the entropy never drops (setup 3 hu / 4 en,
passphrase 5 hu / 6 en). List and count are chosen together so a caller cannot
pair an English list with a Hungarian count. Hungarian is byte-unchanged.

Three claims in the row were wrong and are recorded as such:
  - the RECOVERY CODE is minted by felhom-agent from the EFF list and has
    always been English; the hub does not own it and no row was added.
  - no claim mail states a word count; the only count wording was the bind
    page's passphrase hint, whose English half is now count-free.
  - the proposed phone-safe filter removes 68% of the list (5270 of 7772
    words) and was measured, then declined, with the reason in source.

Also: guide_quote_gate binds the English volunteer guide's three quoted
messages to the controller's English bundle — nothing did, so the guide would
have gone on quoting Hungarian after the fix. Seven decoys, all convicting,
including the name-for-fact one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 07:56:56 +02:00
admin a2c52ebf2a hub v0.118.1: the test mail follows the language too (R-558)
gates / gates (push) Successful in 23s
Found during v0.118.0's own live proof, which is the only reason it was found:
I went to press "send test notification" for the English demo box and read
sendTestEmail first.

It had its own hardcoded Hungarian subject and body and never went through
FormatCustomerEmail, so it was the one customer mail v0.118.0 did not localise
— and it is the only customer mail an operator can trigger on demand, which
makes it the one most likely to be used to check whether the localisation
works. Pressing the button for an English household would have answered that
question wrongly, and convincingly.

The two sentences are extracted byte-for-byte into the bundle, so the Hungarian
test mail is unchanged. Red-proofed against the hardcoded version.

The general form worth keeping: the surface you would use to CHECK a feature is
the one most worth checking first. A broken instrument that reports success is
worse than a broken feature.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-18 16:25:20 +02:00
admin 9167cf53af hub v0.118.0: the household's e-mails follow the household's language (R-558 Part A)
gates / gates (push) Successful in 23s
The hub has written every customer e-mail in Hungarian whatever the box was set
to. The box has published its language since controller v0.247.0; nothing read
it. Now it does.

Nothing an operator reads changes. The Hungarian mails are byte-identical, and
that is a diff rather than a reading: 56 goldens per language captured from
v0.117.0 BEFORE any string moved, and all 56 Hungarian ones pass unchanged after
every sentence was routed through the new bundle.

- internal/i18n: flat bundle, 79 keys, hu authoritative + hu fallback, ceiling 0.
- customerMessages/severityLabels are DERIVED from the bundle, so a sentence is
  written in one place and all 40+ tests that read those maps still work.
- Language order: last reported -> created-with -> hu. reports.language defaults
  to EMPTY, never hu: "never told us" is not "chose Hungarian".
- message_customer on POST /api/v1/event, additive and optional forever, for the
  sentences the box composes and the hub cannot translate.
- The bind page is per-language, and its `expired` state stays Hungarian: it is
  the state an unknown token lands in, so rendering a real English customer's
  token in English would make the LANGUAGE answer what the TEXT refuses to.

Two defects found inside the release:
- R-581: the newest report was picked by received_at, which has SECOND
  granularity, so same-second reports tied and the winner was arbitrary. Ordered
  by the autoincrement id now. GetCustomers() still has the shape - row open.
- R-582: the English copy-guard stems, ported word for word from Hungarian,
  convicted 141 honest sentences. The English claim is a phrase with a modal.

R-555 closed: the language allowlist entry is out of wire_contract_gate.py.
hub_copy_gate.py follows the sentences into the bundle - without that it would
have scanned four files that no longer hold any customer text and reported
success. Three new decoys incl. an innocent control.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-18 16:20:11 +02:00
admin 37ae31fd44 hub v0.117.0: the status follows the configured threshold; slow crash loop and interrupted restore events
gates / gates (push) Successful in 21s
R-549 (operator ruling A): controllerStatus hardcoded 30m/1h while both
staleness checkers and hostStatus read alerting.stale_threshold. Moving the
threshold to 45m would have painted a customer amber 15 minutes before the
alarm could fire - the second definition rollup.go's header forbids. It now
reads the same value, down at 2x. Both 'checker initialized' log lines print
the threshold, which no line did before.

R-539 (ruling 3 of 2026-09-16): controller_slow_crashloop (warning,
operator-only), minted when the agent's slow_crashloop_since moves, with the
fast sibling's first-sight rule.

R-550: restore_interrupted (warning, for the household) allowlisted with a
Hungarian customer message.

Red-proofs, each seen failing then passing: the status test with the old
hardcoded numbers; the checker test with the movement branch removed; the
operator-only test with the registration removed; the household-message test
with the Hungarian entry removed (asserted on the SUBJECT - the body
legitimately repeats the raw message, which my first version of the test
mistook for a fallback).

go build/vet/test ./... green, 18 packages.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 10:20:32 +02:00
admin 3738dfc548 hub v0.116.0: every new customer starts WITH the off-site copy (operator ruling)
gates / gates (push) Successful in 19s
Off-site is ON by default for a new customer — shared, 100 GB soft quota prefilled,
the checkbox kept so an operator can opt a customer out. The reason is this repo's
own [FACT]: the whole-guest tiers do not carry the data drive and a Tier-1 unit has
no file leg, so with this unticked a one-drive box keeps NO copy of the household's
own files. Measured on a fresh box the same day.

The quota is prefilled because the fill warning only fires when quota_gb > 0.

Also registers controller v0.244.0's app_deploy_started / app_deploy_failed in both
allowedEventTypes and customerMessages, per the rule that the two move together.

Red-proofed: dropping the default fails the new render test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 16:56:01 +02:00
admin 926723749d hub v0.115.0: host_* mails skip the quiet hour (ruling 2, R-529); ruling 1 recorded (CC may sign agent_update until the first paying customer)
gates / gates (push) Successful in 23s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 10:53:22 +02:00
admin 07959e61b5 hub v0.114.0: self-bind auto-send while a customer waits for a box (R-509); node_* bypass the quiet hour (ruling 2026-09-15); PBS re-issue adopts an endpoint token (R-511); controller supervisor events (R-523); event registers
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 10:05:41 +02:00
admin 6fd8c87516 doorstep: console is Felhom's (ISO 1.27.0 source), passphrase hand-over copy (hub 0.113.0 source), rulings
gates / gates (push) Successful in 19s
Phase 0: the public ISO never auto-installs by construction (no answer.toml,
G1); the operator re-affirmed the interactive installer 2026-09-14.
- felhom-bootstrap.sh: mask pvebanner.service, write a Hungarian /etc/issue
  (no :8006 admin URL); pairing banner names the Tulajdonosi jelmondat and
  paints through the CONSOLE_DEV seam (R-496). Harness: 8 checks, red first;
  fake hub now sends a pairing code (the banner was never tested, R-502).
- hub: created flash + Credentials block tell the operator to hand the phrase
  over; the self-bind mail names the operator (R-497). Tests red first.
- iso-release-gate G14-G16; domain ruling in 01-topology + CONTEXT; R-494
  narrowed to P3; R-502..R-504 filed; volunteer guide and day-0 A.2 aligned.
ISO_VERSION 1.27.0 (not built, not published).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-14 17:26:53 +02:00
admin f181efd6a7 hub v0.112.0: a floor carries a declared MinAgent past the golden (R-472)
gates / gates (push) Successful in 18s
Operator ruling 2026-09-13. Above the vouched golden, a floor saved with a
declared MinAgent is served under the same agent comparison; an undeclared
one is still held beyond the golden. The declaration is stored beside each
floor as FLOOR=MINAGENT so it never carries to a later floor. Both floor
forms require min_agent above the golden (flash floor_needs_min_agent,
nothing stored). The Hosts page and the API log name the source.

Vouch path and R-120 gate untouched. Scenarios A-E tested; red-proofs A
and C in documentation/audits/rulings-r472-r475-2026-09-13/.
2026-09-13 16:47:11 +02:00
admin db38f4c800 hub v0.111.1: the alarm stops promising a rescue that does not exist, and the arc is closed for beta
gates / gates (push) Successful in 17s
R-434 CLOSED — and the row's own "blocked on R-433" verdict was wrong, which is the point.
The fix is a DELETION, not a replacement: withdraw the promise instead of swapping it for a
new one, and the sentence is true under every possible answer to the provider questions, so
it never needs a second rewrite. A replacement would have been blocked; a withdrawal is not.

  was:  "...still hold the older copy, so this is recoverable file-by-file; it is NOT
         confirmed data loss. Check whether a deletion ran on the box before restoring."
  now:  "...still hold the older copy. The route back out of them is not yet established,
         so treat this as neither confirmed data loss nor confirmed recovery. Get in touch
         before restoring anything, and check whether a deletion ran on the box."

It must not swing the other way either: "your backups are gone" is still usually false.
Clause (a) — the box cannot WRITE into the snapshot area — stands and is re-confirmed.

Tests: offsite_r434_test.go, three, all driving the production path so they assert the
sentence an operator RECEIVES. ASCII-only fragments, positive and negative controls.
RED-PROOF: restoring the v0.111.0 sentence failed all three, on every fragment, with the
offending sentence printed. TestR431_FiresOnAMassDeletion asserted "NOT confirmed data
loss" and caught this fix correctly; its wording fragment is REMOVED rather than updated,
so the wording keeps ONE home.

R-435 written into the detector's own documentation, no threshold changed: it sees a mass
deletion, not one app being wiped (69 across 9 apps -> ~35 needed, one tag is ~9, and
forget --prune groups by host,tags). Says explicitly not to lower the numbers.

THE STOPPING LINE, in all three places — register, 07 section 8 head, STATUS.md.
Deferred set ENUMERATED, not described: 07 section 8 rows 4, 8, 9, 10, 11 (+11b), 12,
each tagged [BETA-DEFERRED]. A number in the brief was wrong and is corrected in place:
six rows are DEFERRED, ELEVEN carry a blank RTO (4,5,8,9,10,11,11b,12,13,14,15); the other
five are blank for reasons that are not deferred work, and row 15 is an open DEFECT (R-104)
that the stopping line does NOT cover. NO STATUS MOVED — nothing was proven today.

Two provider questions drafted, not sent, no API called (11-D stands):
documentation/runbooks/provider-questions-2026-09-01.md, linked from R-95 and R-433, and
tracked by a dated DUE-CHECKS row (2026-09-15) — the 2026-07-27 check that sat unconfirmed
for 36 days is the scar that block exists for.

R-95, R-433 BLOCKED-ON-PROVIDER. R-95's one-day demotion on a clause that did not hold is
recorded; the proposal to rank it back near the top is stated and NOT acted on. R-430 marked
LATENT with its trigger: it becomes live the moment delete is withdrawn, so it is a
precondition on the R-95 build, not a follow-up. The stale ranking paragraph ("armed",
"zero snapshots") is corrected in place, order unchanged.

Register 621 -> 688 lines; 181 rows throughout; open-state 170 -> 169.
No controller or agent change. No golden owed, no floor change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-09-01 18:34:52 +02:00
admin 30681764cb hub v0.111.0: notice a deletion within a day (R-431); correct R-429; re-scope R-95
gates / gates (push) Successful in 17s
THE RECORD WAS TELLING A WORSE STORY THAN THE TRUTH FOR TWO MONTHS, and my own probe is why.

R-429 CORRECTED. Yesterday's spike searched for a directory called `.snapshots`. The vendor documents
the path as /.zfs/snapshot. The probe's CONTROLS were sound and its SUBJECT was wrong, so "not found"
was true and meant nothing. Re-probed at the documented path on both boxes, with controls:

  - /.zfs lists (shares, snapshot) from inside the jail;
  - a write into /.zfs/snapshot is REFUSED - dest open ...: Failure - while the identical write to
    the account home SUCCEEDS and was cleaned up.

That is the append-only property PROVEN rather than cited, and it is the sentence the whole re-scope
rests on. Seven daily snapshots are confirmed in the panel. The mitigation works. What remains true,
and was always the actual finding: the row claiming it had no R-number, its "confirm tomorrow" went
36 days unanswered, and the DUE-CHECKS block built for that class was empty. THE FINDING WAS NEVER THE
SNAPSHOTS - IT WAS THAT NOBODY COULD TELL.

R-95 RE-SCOPED: the box can delete its LIVE repository but cannot write to the daily snapshots of it,
so a deletion costs at most one day plus a per-file recovery - not open-ended loss. The ranking is
Viktor's; it has been #1 since July on the old story.

R-432 FILED: a sub-account sees /.zfs/snapshot EMPTY while the same box holds seven snapshots, so
per-file recovery is operator-only today. One panel read settles whether a NAMED snapshot can still be
entered, which would make it product-reachable.

R-431 SHIPPED. Third signal in OffsiteChecker. On the hub deliberately: a detector on the box is one
the deletion can silence. Threshold REASONED, not invented - over 12 898 reports every decrease lands
on ZERO and predates stats_known, and in the stats_known window there are none, so observed churn gave
nothing to calibrate against. Retention cannot halve a total; a mass deletion goes to ~0. Hence: more
than half, and at least 5. Guarded by StatsKnown (R-331), the declared State (R-204) and run success
(R-100 - whose lesson lives in this very file).

ACCEPTANCE: 9 009 real report points replayed through the detector produced ZERO alarms.

Three red-proofs run. The escalation one only became real after the first version was found HOLLOW -
it re-swept the same report, so the baseline had already moved and the latch was never consulted.

07 row 10's status is NOT moved: the write-refusal is measured, but the recovery ROUTE has never been
walked, which is what PARTIAL means.
2026-09-01 14:25:24 +02:00
admin 1aeaa30c28 hub v0.110.0: allowlist + operator-only for offsite_proof_empty (R-87)
gates / gates (push) Successful in 15s
Controller v0.231.0 adds a nightly job that proves an app's newest off-site snapshot still
CONTAINS that app's data. When it finds one that does not, it emits offsite_proof_empty.

Two register lines, both load-bearing and both in this commit:
- allowedEventTypes - an unallowlisted type is answered 400 and VANISHES, so this entry is
  what makes the alarm exist at all.
- operatorOnlyEvents - a missing customerMessages entry is NOT a routing block
  (FormatCustomerEmail falls back to the raw message, the v0.78.0 defect that register was
  built for). A customer can take no action on a hollow recovery unit.

DELIBERATELY NOT reusing backup_integrity_failed, which is the nearest existing type: it
means THE STORE IS DAMAGED and carries the Hungarian template saying so. Here the store is
sound and the CONTENT is absent - different cause, different action, and telling a customer
their backups are damaged when they are not is the more expensive mistake. Same asymmetry
looksLikeRepositoryDamage is shaped around.

DELIBERATELY no customerMessages entry (the controller's dynamic Hungarian names the app and
what is missing; a template would discard it) and DELIBERATELY not in perAppCooldownEvents
(a fenced act - this job proves ONE app per night, so the coarse hourly cooldown is already
the right grain).

This widens the R-87 task's stated repo scope to felhom.eu/hub/. The reason is recorded in
felhom-controller/CONTEXT.md ruling 4 rather than left as an unexplained diff.

Hub green gate: go build/vet/test all pass, 18 packages.
2026-08-31 20:53:13 +02:00
admin f5c9411e5e R-331 (hub half): the Backup card reads offsite, not the dead backup fields (v0.109.0)
The customer page's Backup card read `Snapshots 0 / Repo Size 0 MB / Integrity
Unknown` for EVERY customer, indefinitely. Measured on demo-hp 2026-08-30 while
that night's controller log said `[offbox] backup OK: 8 app(s) backed up, 67
snapshot(s), 2m14s` and the box held snapshot_count:67, repo_size_bytes:
140829678, stats_known:true.

A card reading "no backups" over a working backup is worse than no card -- the
R-88 direction of failure (degrade to NO BACKUP rather than to UNKNOWN) on the
one screen that answers "is this customer protected?".

The data was never missing. The card rendered the report's `backup` object,
whose snapshot/size/integrity fields have had no producer since slice 8C. The
live numbers are in the `offsite` object, which THIS PACKAGE already reads for
the Offsite page and which monitor.OffsiteChecker already alarms from. Proof the
bytes were arriving: the Offsite page rendered demo-hp's usage as 0.1 GB from
that very object while the Backup card said 0 MB. So this is a render fix over
an existing feed, not a new pipeline.

Not a one-line swap, because snapshot_count:0 means two opposite things --
"holds nothing" and "never measured". R-225 measured that confusion one layer
down. backup_card.go resolves a three-way ruling in Go (a {{if}} chain over
map[string]interface{} float64s cannot keep the absent/zero distinction the card
is entirely about):
  no offsite object   -> "No off-site data reported", and says explicitly that
                         this is NOT the same as "no backups"
  disabled + state    -> names the blocker (needs_credential)
  stats_known:false   -> em-dash + "never been measured". NEVER 0
  stats_known:true    -> the real numbers, INCLUDING a real 0

A pre-v0.225.0 controller sends no stats_known -> false -> "unknown". That
direction is pinned: upgrading the hub ahead of the fleet must not report every
un-upgraded customer as having zero backups.

The Integrity row is DELETED, not re-sourced: nothing produces it, the
controller runs no integrity check, and NotifyIntegrityOK/Failed are called from
nowhere.

RED-PROOF: restore the pre-fix card markup -> all four tests fail, reporting 67
and 134.3 MB absent from the rendered page and the Integrity row present. The
tests drive handleCustomerUnified and grep the HTML on purpose: the defect was
the template's choice of source object, so a test one layer below it would have
been green against the shipped bug.

Green gate clean: 18 packages, rc 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-08-30 18:39:15 +02:00
admin 2fc4a15fa3 R-389: key the operator cooldown per app for app_start_failed; gate 11 makes an unfiled observation refuse the push
gates / gates (push) Successful in 16s
The cooldown key was customerID:eventType plus the tier and run suffixes, and
none of them names an app, so every app going down inside the same hour
collapsed onto one key and only the first was mailed. Measured on demo-hp:
bookstack sent 09:27:51, privatebin suppressed 09:31:51 under
key=demo-hp:app_start_failed.

cooldownStackSuffix is the third sibling of cooldownTierSuffix and
cooldownRunSuffix, and separate for the reason the second one's docstring
already gives: the existing two keep byte-identical semantics for every type
that uses them.

It is ALLOW-LISTED to app_start_failed and takes the event type as well as the
details, unlike its siblings, and that asymmetry is the safety property. The
backup family's cooldown is coarse ON PURPOSE (R-97a, R-182) so one full disk
sends one digest rather than one mail per app - and crossdrive_failed is
severity error, reaches the operator leg, and carries stack_name through a
DIFFERENT struct, so a payload-shape rule would have split it silently. The
hour itself does not change.

Gate 11 refuses a push whose REPORT.md carries an observation with neither
`FILED: R-NNN` nor `NOT-A-FINDING: <reason>`. It deliberately does NOT accept a
passing mention of some other R-number: the lost item cited R-182 as an analogy,
so "cites a register row" would have passed the very item the gate exists to
catch. That discrepancy with the spec is recorded in the gate's docstring.

Registered here and in the controller and agent runners. NOT in the catalog
runner - it has no shared-gate mechanism and appends --all to every gate;
filed as R-391 rather than left as a sentence, which is this session's lesson.

PROMPT-TEMPLATE.md §15.9 corrected: "documented, NOT acted on" was the wording
that invited the gap, and it now names the markers and points at the gate.

R-390 filed for the golden-bake runbook's missing `pveam update`.
Hub tests 709 -> 716.
2026-08-23 13:53:01 +02:00
admin 68a9f5475c hub v0.107.0: the hub rewrote a severity and said nothing (R-387); golden 0.223.0
gates / gates (push) Successful in 17s
One handler, two fields, opposite discipline. An unknown event_type is rejected
with a loud 400. An unknown severity was rewritten to "info" without a word -
and severityNotifies drops "info" before BOTH legs, so the event was stored,
answered 200, and mailed to nobody.

Two shipped features went out that way: DiskAlertKind.Severity emitted "warn"
until controller v0.215.0, app_start_failed until v0.223.0. Measured on the live
hub DB today: 91 app_start_failed events stored all-time, ZERO notification_log
rows before this session - not one, on any channel.

The mechanism built to catch this class was structurally blind to it: the
dispatcher's `unrecognized severity` line cannot execute for anything arriving
over the API, because the coercion one line earlier guarantees the value it
looks for cannot arrive.

The coercion STAYS - a rejected event is a lost event, and losing an alarm is
worse than mis-routing one. Only the silence is fixed: a WARN naming the
customer, the event type and the rejected value.

The dispatcher branch is KEPT, not deleted as dead, and the reason is evidence
rather than caution: cmd/hub/main.go wires dispatcher.ProcessEvent DIRECTLY as
the monitor.EventNotifyFunc for the staleness, host-staleness and offsite-box
checkers, which never pass through the handler. For those it is the only
severity guard there is. All 90 severity literals in internal/monitor are
already valid, so the guard is silent because the producers are correct.

Test count 702 -> 709. Red-proof seen failing: delete the WARN line and the
coercion test fails with "the hub rewrote a severity and said nothing".

Golden 0.223.0 baked and published (sha 9eaf39ac3921...), round-trip HTTP 206.
Vouching is the operator's act and was not done here.
2026-08-23 11:57:26 +02:00
admin ab2262c91c hub v0.106.0: report loss of visibility into the off-site stores (R-339)
gates / gates (push) Successful in 14s
THE GAP, measured not supposed. On 2026-08-18 ep0's PBS proxy was wedged for
9 h 37 m and the hub emitted NOTHING on the operator channel. Both box
checkers hold their last snapshot and return silently on a failed fetch --
correct for a FILL signal, since a missing reading must never be read as 0%,
but it makes a dead off-site endpoint and a healthy one indistinguishable.
The only mails that morning came from the boxes' own backup failures, and
only because the WEEKLY offsite run happened to land inside the window. Two
days earlier nothing would have fired at all.

REACHABILITY is now a second, independent signal on both checkers:
consecutive failed fetch windows, reported past a default 3 windows
(~30-45 min) as pbsdr_box_unreachable / offsite_box_unreachable (warning) on
the customer-less pbsdr-box / pool-box scopes, each with a paired *_recovered
all-clear. Tunable via alerting.box_unreachable_windows (0/invalid -> 3).

THE FILL LOGIC IS UNTOUCHED. No threshold, throttle, band or escalate-once
behaviour changed; a degraded read still drives no transition.

Three decisions a later reader would otherwise "fix" back, so each is
argued in-code:
  - the unreachable event REPEATS rather than escalating once. The band shape
    would give exactly ONE mail at ~minute 30 of a nine-hour outage, and one
    mail is missable. It leans on the dispatcher's 1 h operator cooldown to
    become an hourly "still blind" heartbeat.
  - ErrUsageUnsupported is NOT blindness: an old ep0 answers "no such op",
    which means we reached it. Counting it would alert for days on a healthy
    pre-update endpoint.
  - born-blind is reported: the counter is not gated on having a snapshot, so
    a hub restarted INTO an outage still speaks. last_ok is OMITTED rather
    than zero-valued -- a fabricated timestamp reads as "it was fine until
    then".

Both recoveries are severity "info" and severityNotifies drops "info", so
they are registered in recoveredPairedDownTypes or the operator hears that
the tier broke and never that it healed. A cross-package test drives
ProcessEvent and asserts an actual operator MAIL, not a map entry -- a green
checker test proves nothing about the seam (agent v0.91.0 shipped fully green
with SetAuthSink never called).

Tests: box_reachability_test.go (Scenarios A-F) + dispatcher_box_reachability
_test.go (wiring). Three red-proofs run and reverted, each seen failing with a
message naming the right cause: threshold 3->1, the sentinel counter guard,
the pairing entry.

Register: R-339 filed and marked SHIPPED (PROVEN-LIVE still owed -- no real or
constructed outage has exercised the emit path, and one cannot be manufactured
against Tier-2 ep0). R-340 filed: the reachability read rides ep0's LOCAL API
daemon, which the incident explicitly cleared, so this check would have shown
GREEN for all 9 h 37 m -- the honest boundary, recorded rather than glossed.
R-336's next-step corrected: pvestatd's interval is NOT tunable (Proxmox staff
have said so); the only lever is disabling the storage entry, which collides
with the agent's consume-the-one-time-secret path. Doc-only, no agent code
touched.
2026-08-18 19:27:33 +02:00
admin b03a105375 hub v0.105.0: the third name, a machine told to be quiet, and a guard for the hub's own words
gates / gates (push) Successful in 17s
Hub only. No controller change, no agent change, no wire change — nothing to bake.
demo-hp untouched: the operator is re-deploying it this evening.

R-323 — the five-word phrase is „Tulajdonosi jelmondat". It was „Visszaállító
jelszó": one word from the name retired last week, and false besides — it restores
nothing, it proves the account owns the box being bound. Five sites, all in the hub;
felhom-controller and felhom-agent carry the name nowhere, so no halt and no bake.
Both suggested names were rejected with reasons: „Fiókjelszó" would collide with the
dashboard login (a DIFFERENT real secret), and „Összekötési jelszó" would leave the
two factors on this page separated only by kód-versus-jelszó — the exact shape being
removed, since the other factor is the „Párosító kód". The chosen name differs on
both axes, stem and noun. Naming only; the acceptance pin drives the real handler.

R-324 — the hub's customer copy is under a guard for the first time. Retired names
banned across all 95 hub files; retrieval stems registered in four declared customer
surfaces. The selftest found a defect in its own instrument on the first run. One
shared vocabulary in scripts/, drift-checked into the controller gate rather than
copied (R-325 removes the scaffold).

R-321 — a machine we told to be quiet is no longer reported as dead, and it was two
doors, not one: because the state is RECORDED rather than deleted, the morning
deadline check can skip it too. A deleted state returns "", which is not "down" —
R-195's shape returning through a second door. The clock runs from the report the hub
can see, so re-enabling starts it there and emits no recovery for an outage that never
happened. Three red-proofs; the one that matters showed a genuinely dead machine
sitting at "disabled" when the suppression was made unconditional.

R-326 — "which claims are unproven" is answerable by a command now. The nine I have
been repeating was the count of claims the 9 August pass DOWNGRADED, not the count of
unproven ones. The real figures: 55 claims, 23 walked, 32 not — and only 6 of those 32
cite evidence. Its first run found a stale claim (R-327).
2026-08-13 15:50:32 +02:00
admin 4d6ec7c7bb hub v0.104.0: the guest network gets a reader (R-319), and the hub half of the naming (R-295)
gates / gates (push) Successful in 14s
Four paper debts and one fact given a reader. Hub-only — nothing to bake.

A4 — the entry about "the tester's machine" named a risk correctly and labelled it
in a way that invited deleting it. Established from the hub's own store: `peti-felhom`
is a REAL machine (482 reports, 2026-02-27 → 2026-07-15, a named person's own box) and
the 3.6 GB with no key and no backup is real. `david` → `tester-1` is a DIFFERENT record
with no host, no escrow and no report, ever — deleted 07:55:49 and re-created 07:56:47
this morning. The prompt's premise conflated the two; the register now says which is which.

A1 — R-312/R-313/R-303 recorded as DECIDED with their re-open triggers, and moved out
of STATUS's "Waiting on you", which is now empty.

A3 — day0-install §C.1 said pushing the installer publishes it. It has not since
R-110. Corrected, with the two manifest pins named and an outside-verification command;
the one copy that repeated it (a dated audit, true when written) carries a superseded note.

A5 — standing rule 5: evidence comes off the machine at the end of the phase that
produced it, before any revert. Earned twice in three days on the same box at the same
point (R-320). Four homes, plus what to do when it is already gone.

R-295 hub half — „Beállító kód" everywhere; „Visszaállító kód" retired. New `reenroll`
mail kind so the mail names the page a REBUILT box actually shows („A szerver
beállítása"), not the „Elfelejtett jelszó" page it has no login screen to reach.
Naming only; the acceptance pin proves the secret is untouched.

R-319 — the hub models `guest_net` after 23 days of receiving and discarding it. The
signal is `heals_last_hour`, not `state`: a guest the watchdog keeps repairing reads
healthy between repairs. `heal_succeeded` decoded too (R-260's lesson). Unknown is never
drawn as healthy — three absences, three sentences. No alarm, deliberately.
Three red-proofs, mutations asserted applied. Wire-gate checked tags 182 → 190.

B1 — the operator's 2026-08-12 dispositions were NOT in the register; they are now.
Third allowlist kind for the five ruled "no reader wanted"; `reporting_disabled`
reclassified redundant. 8 read · 5 deliberately unread · 1 redundant · 6 still owed.

Also filed: R-321 (a deliberately-silent box still alarms stale/down — the checker is
age-only, and decoding the flag would not have fixed it), R-322 (the claim guard has
never scanned the hub; a hand scan returns zero, so it is a scope gap, not a defect).
2026-08-13 10:50:12 +02:00
admin 6362bb6cb6 hub v0.103.0 — a host can read the packages we kept for it (R-311)
ListSupersededEscrow had zero production callers for nineteen days. It is the only
reader of a retained identity_blob, so the retention shipped in v0.93.0 was
material the product could not reach - proven on the fixture 2026-08-12, where a
code that opens a retained package was answered as a code that opened nothing.

New GET /api/v1/hosts/<id>/escrow/retained: self-scoped exactly as the current-row
GET, same recovery-mode gate, same audit event written BEFORE the bytes leave,
capped at 16.

Rows with a NULL identity_blob are WITHHELD and returned as unopenable_count.
They retain the PBS key, not the repository password, so they can never open what
the caller is asking about; serving them would have the agent try packages that
cannot succeed and would let the screen claim an earlier package is openable on
exactly the boxes the original defect hurt. The count is returned because their
existence is load-bearing and underivable.

The trade, stated rather than waved through: the hub still cannot read any of it -
sealed bytes in, sealed bytes out, no decrypt path, no recovery code ever held.
What widens is volume, bounded by self-scope, the recovery-mode gate and the cap.

The response is a NAMED TYPE, not a map, so the wire-contract gate can resolve it;
the wire is declared as a fourth ROOT and the gate now checks 182 tags rather than
174. A positive control shows that check is name-presence, not decodability -
filed as R-315 rather than reported as coverage.

Six tests through the real endpoint; four red-proofs asserted applied.
2026-08-12 18:43:07 +02:00
admin b55fc17d82 hub v0.102.0 — refuse to vouch a version that cannot be installed (R-273)
The guard owed since Friday morning. Agent v0.128.0 was published as a package
and never git-tagged; it was vouched here; and because felhom-host-install.sh
fetches an agent's configs from raw/tag/v<version>/configs/, every fresh install
and reinstall died at step 5 of 8, as root, on a virgin machine, for most of a
day. handleSetArtifacts is the sole UI path to SetArtifactManifest, so the check
belongs here and nowhere else.

TWO LEGS, because both failed inside two days: the TAG (missing, R-273) and the
PACKAGE (pruned from under a still-tagged version, R-287). Either alone catches
one of them.

It asserts configs/felhom-mkfs-guarded.sh -- the FIRST of the installer's sixteen
fetch_raw calls and literally the file whose 404 broke Friday. A test pins the
constant, because probing a path that merely exists is how it stayed invisible.
The golden gets the package leg only: it has no config tree, so a tag probe would
assert something the installer never does.

"Could not verify" refuses too, with its own message. No override -- the registry
is the operator's own server, so if it is unreachable the vouch can wait.

ORDERING IS LOAD-BEARING AND A FAILING TEST FOUND IT. The probes run before
resolveArtifactSHA, whose flash conflates "missing", "unreachable" and "bad sha".
Probing first means an unreachable registry is reported as unreachable.

Five scenarios each naming the wrong outcome; three red-proofs, mutations asserted
applied and reverted. With the tag check removed, scenario A reports artifacts_set
-- Friday's exact defect returns.
2026-08-09 19:13:45 +02:00
admin 7264f02172 hub v0.101.0 — memoise the artifact dropdown for 60s (R-267, operator ruling)
gates / gates (push) Successful in 28s
v0.100.x removed the serialisation: 26.2s -> ~9.85s mean. What remained was one Gitea package SEARCH
per dropdown at 0.20-3.8s depending on load, which concurrency cannot help.

Memoised for 60s IN MEMORY. The TTL was ruled by the operator against the workflow that cares: a
bake-and-vouch session publishes an artifact and comes straight here to select it, so a minute is
short enough not to be noticed and long enough that every reload in that session is instant.

NOT persisted. Gitea IS the store for both the version list and the sha; a copy in hub_settings would
be a second source of truth that can drift from the registry it describes, and the operator reads the
sha here to confirm what they are about to vouch. An in-memory cache dies with the process and can
never be mistaken for a record.

A failed resolve is NOT cached — a blip must not pin an empty dropdown for a minute. But an empty
list from a package that genuinely has no versions IS cached, because 'we found nothing' and 'we
could not look' are different answers (CONTEXT S-39, applied to a list instead of a figure).

THE FIRST VERSION OF THIS GOT THAT WRONG: the comment said only successful resolves were cached and
the code cached the empty list anyway. TestArtifactChoices_FailureIsNotCached caught it before it
shipped — which is the argument for writing the test that asserts the comment, and the same class
this session spent the day closing.

go build/vet/test green, go test -race clean, run separately from this commit.
2026-08-08 20:08:51 +02:00
admin e348c4ef6e hub: the Gitea client keeps its connections (MaxIdleConnsPerHost was 2)
gates / gates (push) Successful in 28s
Third and last leg, found the same way as the second — by not accepting that the numbers matched the
arithmetic when they did not. After the fan-out and the side-by-side resolve the page was ~11.9s mean
where ~3s was predicted.

Cause: the client used http.DefaultTransport, whose MaxIdleConnsPerHost is 2. Above that Go opens a
connection per request and discards it after, so under a 16-way fan-out almost every call paid a
fresh TCP setup AND a fresh authentication. Authentication is the expensive half: unauthenticated
/api/v1/version answers in ~0.03s while an authenticated package call takes ~0.24s against the same
Gitea instance.

Transport sized to the fan-out: MaxIdleConnsPerHost 16, MaxConnsPerHost 16 as a ceiling so a large
package list can never stampede Gitea harder than the fan-out needs, IdleConnTimeout 90s.
2026-08-08 17:55:18 +02:00
admin 9ce8c631d3 hub: resolve the two artifact dropdowns side by side, not one after the other
gates / gates (push) Successful in 30s
Follow-up to the fan-out, and the reason for it is worth recording: THE FIRST FIX DID LESS THAN THE
ARITHMETIC PREDICTED. Concurrency took the page from 26.2s to ~11-18s, not the ~2s expected, so the
gap was chased instead of declared closed.

What it found: the slowest single call on the page is not a per-version sha lookup at all, it is the
PACKAGE SEARCH (/api/v1/packages/admin?type=generic&q=...), measured in-cluster at 1.1-2.2s each
against ~0.24s for a file's metadata. Per-version fan-out cannot touch it — there is one search per
package and they ran in series.

The two dropdowns are independent, so they now resolve side by side, overlapping both searches and
both fan-outs. go test -race clean on the new concurrent paths.

Gitea's latency on this box is load-dependent and varies 2-4x between samples, so the CHANGELOG
quotes a range rather than a single pair of numbers.
2026-08-08 17:50:13 +02:00
admin 7855d6355c hub v0.100.0 — the Configuration page took 26 seconds, and it was never hashing anything
gates / gates (push) Successful in 21s
MEASURED, NOT GUESSED: GET /configuration -> HTTP 200 in 26.2s.

The reasonable guess was that it hashes the artifacts on page load. It does not, and the code already
said so: Gitea stores each package file's sha256 and gitea.FileSHA256 reads it as metadata — "a cheap
metadata call, the artifact bytes are never downloaded". The cost was never CPU.

IT WAS LATENCY x COUNT. artifactChoices made ONE SERIAL round-trip per version, for two packages,
capped at 20 each: 2 x (1 version list + 20 sha lookups) = 42 sequential requests at ~0.6s each out
through the public ingress. 42 x 0.6 = 26s, which is what the clock said.

1. The sha lookups now run CONCURRENTLY, bounded at 8 in flight. Order preserved by writing into a
   slot rather than appending — the dropdown is newest-first, and a scrambled sha would show the
   operator a hash belonging to a DIFFERENT artifact. A failed lookup still drops that version only.
2. The client talks to Gitea IN-CLUSTER (http://gitea.gitea-system.svc.cluster.local:3000,
   overridable via GITEA_API_URL). Measured from the hub pod: 0.11s against 0.26-1.16s, because the
   public path adds DNS, the ingress hop and a TLS handshake to each of the 42. Plain HTTP is safe
   ONLY because it never leaves the cluster network — the registry token rides the Authorization
   header, so this must not point at a public host without TLS. Unreachable -> the existing graceful
   degradation to manual text entry, unchanged.

DELIBERATELY NOT DONE: caching the sha in the hub's own database. That was the other half of the
proposal and it is the wrong shape. Gitea already IS the store; a copy in hub_settings would be a
second source of truth that can drift from the registry it describes — and the operator reads exactly
this value to confirm what they are about to vouch, so a stale one would be a confident wrong answer.
The same reasoning golden_currency_gate.py already records for the vouched version. With the fan-out,
a cold load needs no cache to be fast.

The cap stays at 20 and now bounds the FAN-OUT too, not just the rendered list.

Tests pin order (and that each sha belongs to its own version), per-version failure isolation, and
THE CONCURRENCY ITSELF — a wall-clock assertion plus an in-flight counter, so a fast run cannot be
luck, and an upper bound so a large package list cannot stampede Gitea. Red-proof: reverting to the
serial loop takes 861ms where the concurrent one takes 150ms, and the test fails naming the
26-second page.

go build / go vet / go test ./... green (18 packages), run separately from this commit.
2026-08-08 17:41:52 +02:00
admin b080ecf411 hub v0.99.0 — the hub can see whether the operator can get in (R-260); G-1 gate closes, R-247 closes
oobDegraded tested five things and the sixth never arrived.

The agent has emitted `operator_key_configured` on every heartbeat since v0.72.0 — the SAME version
that introduced the `oob` stanza carrying it — and store.HostOOBRow mirrored five of the agent's
eight OOB fields. With no field for it, encoding/json discarded the fact on arrival, so a box with
felhom-sshd active, reachable, a valid config and a configured peer reported `ok` with NO OPERATOR
KEY INSTALLED AT ALL. Not a wrong answer: an answer to a question nobody was asking.
`operator_peer_configured`, which the hub did read, only says the peer IP is in desired-state — that
OOB is MEANT to work, not that entry is possible.

Now decoded: operator_key_configured, plus wg_handshake_age_s and healed_at. The last two ride the
ALERT TEXT and are deliberately NOT in the predicate — widening a check beyond the fact that is now
arriving is how a check stops being read.

SCENARIO F, decided on a measurement rather than a preference. operator_key_configured decodes as a
POINTER: nil = the agent never said, reported distinctly and never as ok. The version gate was
rejected because the field and its stanza shipped in the SAME agent version (v0.72.0), so a stanza
without the field cannot come from any released agent; the fleet is 0.113.0/0.127.0 and the vouched
floor is 0.127.0. Handled explicitly anyway and pinned, because "cannot happen" is a claim this
project has been burned by.

THE MESSAGE NAMES THE FAULT. oobDegradedReason is the single source for both predicate and text, so
the alert can never name a different fault from the one that fired. The old form derived it
separately and had a vocabulary of two — unreachable, or config invalid — with no way to say the key
is missing. The operator reads this at 07:00.

TESTS DRIVE THE DECODE BOUNDARY. Every hub OOB test before this built a HostOOBRow by hand, and a
test written that way CANNOT SEE A FIELD THAT NEVER DECODES — which is how this held a green suite
for five weeks. The pre-existing fixture oobReport() also omitted the field, so those scenarios ran
against a report shape no released agent produces (same family as R-262). Both fixed.

Red-proofs, 8 expected outcomes and 0 wrong, each with the mutation asserted applied: dropping the
field returns the false ok; an unconditional check alerts a healthy box; unknown-as-ok restores the
silent pass.

G-1 CLOSED — scripts/wire_contract_gate.py shipped as ranked, built BEFORE the fixes and seen
failing on 40 fields (documentation/tests/wire-contract-gate-2026-08-08/BEFORE.md). Two instrument
defects the control caught first: a substring false negative (grep -F healed_at matched
privsep_healed_at) and treating dr_recipe as wholly opaque when its top-level sections ARE decoded
through an allow-list that already cost offsite_restic (R-122).

The prompt for this session said "465 emitted tags, eight unreachable". Checked against the repo:
R-260 said "at least eight DECISION-BEARING facts", never eight tags. The real count is 40.

R-260 CLOSED (class gated, sharpest instance fixed). R-247 CLOSED (controller v0.209.0). R-264
MINTED and OPEN — the 21 facts with no consumer, allowlisted with reasons so that gating the class
could not be mistaken for deciding them. Still open and named: R-246, R-255..R-259, R-261..R-263,
and C7's test-comment half.

Capability map checked: it claims OOB access is implemented, never monitored, so no row was untrue;
what was untrue sat one layer down and the row now records it.

repo_gates --fast: all 8 OK. go build/vet/test green in hub, run separately from this commit.
2026-08-08 08:47:02 +02:00
admin ac4b2a4ba9 hub: drop the retained recovery package when the box declares its set-aside history deleted (R-241)
The hub half of the controller's abandonment countdown, and the ONLY reason
felhom.eu was touched for R-241 at all.

A customer who abandons their old off-site history gets a 14-day countdown. At
the end of it the controller deletes the set-aside restic store and then
DECLARES offsite.abandon_purge_requested in its report until the retained
sealed package that protected that store is gone too. Removing only one half
leaves a state that asks a question nobody can answer: a package that opens
nothing, or ciphertext nobody can ever decrypt.

store.PurgeSupersededEscrowForCustomer is the one place R-198's retention is
ever undone, and its doc comment says why that is legitimate here. It NEVER
touches host_escrow - the current package covers the key the box is using now
and is what makes its live backups recoverable. Only host_escrow_superseded
rows go.

The handler acts on the box's DECLARATION, never an inference, on the same
principle as offsite.state: the hub cannot see that a remote store was deleted
and the box can.

It is placed immediately BEFORE the ACK is built, deliberately.
GetEscrowStatusForCustomer is read after it runs, so the SAME response that
carries the request's effect is what closes the box's two-phase commit - no
second round-trip, and no window in which the box believes it is still owed.
The declaration repeats on every report until that ACK stops reporting a
superseded package, so a lost request retries by itself.

A purge failure is logged at ERROR and never swallowed: the box keeps
declaring, so it retries, but an operator must be able to see that the two
halves are apart right now. An idempotent re-declaration (already purged, the
box has not yet seen the confirming ACK) logs at DEBUG and is not an error.

Audit event offsite_abandon_purged is hub-internal, like the pbsdr_* and
offsite_selfheal_* events - allowedEventTypes governs the box-pushed
POST /event surface, not this.

No agent change. No deletion has been performed against any real store.

Green: go build, go vet, go test ./... all clean in hub/; repo gates OK.
2026-08-07 11:48:10 +02:00
admin a7f1d277b1 hub: the held-floor REASON must match the CAUSE (CAMPAIGN-11 follow-on)
gates / gates (push) Successful in 7s
v0.97.0 introduced a second hold reason and left both surfaces printing the first. The
freshly deployed hub logged, for the campaign box:

  managed floor HELD for c11: agent "0.125.0" < MinAgent 0.113.0

which is FALSE — 0.125.0 is ABOVE 0.113.0. That box is held because its floor sits above
the vouched golden, not because of its agent. CLAUDE.md's corollary exactly: when a verdict
changes which field it counts from, the alarm text has to change with it, or a true alarm
reads as one to dismiss.

Both the ACK log line and the Hosts-dashboard HeldReason now come from one
ManagedFloorDecision.HoldReason(), and TestResolveManagedFloor_HoldReasonMatchesTheCause
pins each reason to its cause.
2026-08-05 17:53:24 +02:00
admin 7e1d2898bd hub v0.97.0 — the floor stops being served past the agent it depends on (CAMPAIGN-11)
gates / gates (push) Successful in 8s
R-216, the hub half. ResolveManagedFloor's own comment says it exists to "never push a
controller past the agent it depends on", and it compared against ArtifactManifest.MinAgent
— which by ITS own comment describes the GOLDEN's controller. publish-train-rules.md rule 3
states the rule about the FLOOR's controller. Measured live: golden 0.192.0 / MinAgent
0.113.0, floor 0.200.0, agent 0.120.0 — served, and the box was pushed onto a controller
needing agent 0.125.0.

A floor ABOVE the vouched golden is now HELD with its own reason (HeldBeyondGolden), reusing
Part D's dashboard visibility. Nobody types a number twice: the vouched MinAgent keeps its
meaning, the guard stops applying it to versions it does not describe. An uncoupled release
is untouched; an unparseable golden degrades rather than gating.

R-222: the report ACK's escrow object gains superseded_present / superseded_at, counting only
rows that actually carry an identity blob. One boolean and one timestamp, for one message.
No read path — that link is still unbuilt.

Red-proof: removing the floor-above-golden branch reproduces the campaign's measurement.
2026-08-05 17:49:04 +02:00
admin f62a115891 R-204 item 4 (hub half): the hub answers a rebuilt box's request (hub v0.96.0)
New internal/offsiteheal, the sibling of pbsdrheal: it acts ONLY on the state the
box declares, sustained across two distinct reports, re-staging the stored
credential before ever minting a new one. A healthy box is a pure no-op; it never
blind-timer-reissues and never re-runs a provisioning step.

RESTAGE IS POSSIBLE because the stored value survives a consume — established from
the schema and ConsumeOneTimeSecret (which stamps consumed_at and nothing else),
not inherited from the PBS analogy, and pinned by a test that asserts the SAME
value comes back.

reportHasOffsite is TIGHTENED to require enabled:true. Its comment asserted that
presence == applied-on-the-box, and the declaration deliberately breaks that
premise; left alone it would have read a request for help as proof the tier was
applied. Provably a no-op for every report shape that existed before, because an
attached object has always carried enabled:true.

R-192's guard half is CLOSED BY REPLACEMENT: the delivery checker's counting
inference read the OLDEST 500 reports after a consume — all predating a rebuild,
which is why demo-hp sat stranded for 108 reports under a confident regressed-shape
verdict. A declaration outranks both inferred shapes, and the checker stands down
with a record so the two mechanisms cannot double-issue.

No escrow ceremony is ever run or requested: credential automatic, key
customer-present.
2026-08-05 10:48:18 +02:00
admin d1a8edb332 R-196 / R-204 item 2: a re-issue no longer marks a healthy escrow stale (hub v0.95.0)
ReissueCredentials marked the escrow stale on every re-issue, on precautionary
grounds — the box's re-apply MIGHT mint a fresh repository password. It usually
does not. A stale flag withholds restic_pw_sha256 from the ACK, which stops the
controller's auto-confirm, which leaves EscrowState pending, which makes
OffboxRunnable false: every off-site backup refused on a box whose key was never
in doubt — and the customer told to re-run the one ceremony that would have
superseded the key just recovered.

The case it guessed at is measured elsewhere: the controller's Scenario-F
re-check compares the sealed hash against the live repo password on every ACK
(and the mark was BLINDING it by emptying that hash), and R-197's
offsite_repo_key_changed fires on a proven difference across a supersession.

offsite_reissued is unchanged. MarkEscrowStale is kept without a caller so a
future EVIDENTIAL writer has the mechanism, with a test pinning it live.
TestReissue_InvalidatesEscrow is replaced by its exact inverse.
2026-08-05 07:17:29 +02:00
admin 435f4a5229 hub v0.94.0: a box can fetch its own sealed recovery package (R-199 link 6)
gates / gates (push) Successful in 7s
Link 6 of the recovery chain had no client. The hub has served the identity blob since
slice 10D from handleReEnroll / handleGetRestoreDirective, gated on operator-armed recovery
mode and the global key -- and nothing in the agent, the hub UI, any script or any runbook
ever called either. The only documented retrieval was sqlite3 writefile() by hand on a
kubectl cp-ed database.

GET /api/v1/hosts/{host_id}/escrow is the box-authenticated mirror of the PUT that put the
blob there. Self-scoped (a per-host key reads only its own; global may read any). A host with
no bundle gets 200 {present:false} -- a 404 is indistinguishable from an unknown host and a
bare empty 200 from a zero-length blob.

THE TRADE IS RECORDED IN THE HANDLER, not inferred: obtaining the blob used to require the
operator to arm recovery mode; now whoever controls a rebuilt box can obtain it with that
box's own credential. They still cannot open it -- the hub has never held R and a wrong code
fails closed at age's scrypt KDF. The mitigation is that every retrieval raises
escrow_blob_served (warning, operator-only), recorded before the bytes leave.

escrowSelfServiceRetrieval is the single decision point: flip it to false and the endpoint
additionally requires recovery mode, changing nothing else.

The operator-driven DR path is untouched, pinned by a test. Red-proofs observed: removing the
ownership check serves host B's blob to host A; removing the record makes it silent.
2026-08-04 13:39:27 +02:00
admin 91cabdde1b hub v0.93.0: the retention keeps the key it was built to keep (R-198) + three honesty fixes (R-197, R-192, R-196)
gates / gates (push) Successful in 7s
R-198 — host_escrow_superseded shipped with `blob` (the K-escrow / PBS datastore key) and
identity_blob was added to host_escrow LATER, never here. The offsite restic REPOSITORY
password lives in identity_blob. So demoteCurrentEscrowTx -- whose own comment calls it "THE
ONE escrow row-copy routine" -- retained the whole-guest key and silently dropped the off-site
data key, which is the secret the retention was built to preserve. And because the copy happens
as the new blob overwrites the old, the destroying act was the ESCROW CEREMONY: the exact thing
a rebuilt box tells its customer to run, on a card promising in Hungarian that the old backups
stay recoverable. Both demo boxes crossed that line on 2026-08-04.

  - identity_blob added to the table (CREATE + additive ALTER) and carried in the shared copy
    routine, so BOTH callers are fixed at once: re-escrow and host-delete demotion.
  - ListSupersededEscrow reads it back; store.HostEscrow gains IdentityBlob.
  - CountCurrentEscrowWithIdentity is the census of who the fix protects.
  - Nothing is backfillable: pre-v0.93.0 retained rows have no blob and their sources are gone.
  - Tests assert the CONSEQUENCE (a retained row can still yield a repo password), which is why
    the pre-existing retention test stayed green for two months asserting the mechanism.

R-197 — SaveHostEscrow returns the hash it replaced; the escrow PUT raises
offsite_repo_key_changed (warning, operator-only, edge-triggered) when both hashes are known and
differ. No hash value travels. Severity chosen for the world v0.93.0 creates: with the identity
blob retained, a changed key is "this history now depends on an older recovery code", not a loss.

R-192 (half) — the stuck alert now reports the two shapes it actually covers, burned and
regressed, each stating its own measurement; the regressed text withdraws the Re-issue
recommendation. Every self-heal refusal leaves a notification_log row with its reason. The
guard's logic is unchanged; its 500-oldest-reports scoping stays OPEN and the window is named in
the alert text so the limitation travels with the number. offsite_delivery_stuck and
offsite_credential_restaged are added to operatorOnlyEvents -- neither was registered and neither
has a customerMessages entry, which is not a block.

R-196 — five comments (not the three the spec expected) claimed ReissueCredentials rotates the
restic repo password. It resets the PROVIDER password and cannot touch the repo password, which
is generated on the box. All five corrected; the staleness mark documented as precautionary. The
BEHAVIOUR stays open.

Not in this release: R-199, R-200, R-201 remain open -- the chain that hands the key back is
still unassembled. Part 5 hit its gate; the orphan card is untouched (R-202).
2026-08-04 12:56:58 +02:00
admin 7fff45d688 R-195: a customer with no machine ever bound does not alarm (hub v0.92.0) + R-193/R-192 spike
gates / gates (push) Successful in 7s
Part 4 (ships): `david` — a prospective customer with hosts=0, host_deletions=0,
reports=0 — e-mailed an expected_dbdump_missed ERROR at 03:00 UTC three mornings
running. The existing down-skip could never cover it: it reads the staleness
checker's state, which is seeded from a query over the `reports` table, so a
customer that never reported has no state at all and GetState() returns "" rather
than "down". store.HasEverBoundHost (hosts row OR host_deletions tombstone) is
consulted once per customer at the top of the deadline loop. The discriminator is
"was a host EVER bound", never "has a report arrived" — a box installed and never
heard from is a real fault and keeps alarming. Fail-OPEN on a read error. Red-proof
observed: removing the guard fails with `got [expected_dbdump_missed]`, verbatim the
event david sent.

Parts 0-3 (spike, NO production code for R-193/R-192):
audits/SPIKE-offsite-credential-recovery-2026-08-04.md establishes that the one-shot
provider password is the RECOVERABLE secret and the restic repository password is the
irreplaceable one — and that a guest rebuild mints a fresh one, orphaning the previous
off-site history. Measured without touching a box, by comparing
host_escrow.restic_pw_sha256 against host_escrow_superseded: BOTH demo boxes changed
(demo-hp 15 snapshots / 40.9 MB, demo-felhom 36 snapshots / 1.14 GB). demo-felhom's
"lucky" 76-second recovery restored delivery and not the repository, silently, for 13h.
ReissueCredentials does NOT rotate the restic password (R-39's record and two hub
comments are wrong -> R-196); candidate (b) is not implementable against a
zero-knowledge escrow; candidate (a) already exists as F3 and is wired to the wrong
event. Ends in ranked options and an unanswered question for the operator.

R-195 SHIPPED; R-196 + R-197 filed; R-192 + R-193 updated, neither closed.
2026-08-04 11:04:39 +02:00
admin 046df303b6 hub v0.91.1 — observation may only WIDEN a tier's window, never tighten it (R-86)
gates / gates (push) Successful in 7s
Found by checking v0.91.0 against the live box, not by review. demo-felhom's two
retained PBS snapshots sit 8h54m apart (one is a healing artefact), so the
mean-gap estimator reads a WEEKLY tier as nine-hourly: x4 = 36h, the 7-day floor
lifts it to 168h, and a weekly tier proved weekly reaches ~8.25d of proof age.
The false alarm this task exists to prevent would have returned within a week, on
the box it had just shipped to.

restoreProvenWindow now takes max(observed, declared). A gap SHORTER than the
declared rhythm is routine and means nothing (a retry, a manual run, a heal, a
catch-up); a gap LONGER than it is real information. Cost stated: a tier running
faster than its declared rhythm gets a slower stale signal — the right direction
for a signal that means 'unverified', since 'broken now' is a different event.
2026-08-03 15:16:59 +02:00
admin 323f45a5ef hub v0.91.0 — the staleness window learns each tier's own rhythm (R-86 Part 2)
gates / gates (push) Successful in 7s
Ships WITH agent v0.121.0, not after it. The agent now proves a tier once per
ARCHIVE GENERATION, so a weekly tier is proved weekly — in perfect health. The
flat 7-day restoreProvenStaleAfter derived its number from the 24h cadence R-86
removes, and a healthy weekly tier's proof age reaches EXACTLY 168h just before
its next proof: it sat ON the line, so any ordinary delay tipped it into a
nightly alarm about a working system.

restoreProvenWindow(tier, observed, ok):
- the tier's own archive interval, OBSERVED from reports the hub already holds
  (pbs_snapshots + successful backups attributed by TARGET TYPE, slice A.4)
- x4 generations = the same tolerance the flat constant expressed
- floored at 7d (never tighter than before), capped at 12d (strictly inside the
  2-week offsite retention)
- falls back to the DECLARED rhythm (26h host / 8d offsite — the thresholds the
  backup-freshness checker already uses) when history is too short to observe
  one; falling back to the FLOOR would recreate the false alarm on a fresh box

Kept: absence is UNKNOWN until the anchored window passes; the signal stays
edge-triggered; failed and stale remain distinct events. Every reason string now
states the window it was judged against (R-100's corollary).

Also backfills the missing v0.90.1 CHANGELOG entry (deployed since f21e7ca), and
records the operator's 2026-08-03 ruling that ep0 is Tier 2 / protected.
2026-08-03 15:03:35 +02:00
admin f21e7caed1 hub v0.90.1 — the digest's per-app lines stop repeating the filesystem figures (R-182)
gates / gates (push) Successful in 7s
Found by reading the first REAL digest, not by design. Every app row ended with
the same usage clause the mail already prints once on its own Filesystem line.
On a two-app box that is untidy; down a list of a dozen it is the same forty
characters twelve times, pushing the part that DIFFERS off a phone screen at
07:00 — the only moment this mail has to work.

The reserve's refusal message is authored for a single-app alert where naming
the filesystem is right, so the message is unchanged; the digest trims the
duplicate when rendering. trimRepeatedUsage removes ONLY an exact
"— <target path>:" suffix, so an unrelated reason is untouched and a reason that
is nothing but the usage clause is left alone rather than emptied.

Also updates TestRecoveryUnitCaptureFailed_NeverReachesTheCustomer, which
required the OPERATOR to be emailed a per-app capture failure. That was correct
when the event was the only signal and is wrong now that it is the record and
the digest is the notification. Its customer-safety claim is unchanged and is
why the test still exists; the operator assertion is inverted with the reasoning
written in place, and R-158's guarantee is shown to have MOVED, not weakened.
2026-08-03 13:54:02 +02:00
admin dd40f85bb8 hub v0.90.0 — a dropped notification leaves a trace, and the backup digest arrives (R-182)
gates / gates (push) Successful in 7s
processOperator's cooldown no longer returns bare. It dropped the event BEFORE
LogNotification, so a suppressed operator alert and an event that never happened
were indistinguishable — from the operator's side and from the hub's own records.
Measured 2026-08-03: nine recovery_unit_capture_failed events arrived, two were
mailed, seven left no row anywhere. That is why the defect took a day to get the
right way round: there was nothing to read.

A suppressed operator event now writes a `suppressed` row carrying the message
and the key that suppressed it. This applies to EVERY operator event, not only
the one that exposed it. It does NOT change the cooldown's duration or semantics.

backup_run_failures: the per-run digest. In allowedEventTypes AND in
operatorOnlyEvents — allowlisting alone does not make an event operator-only,
and FormatCustomerEmail falls back to the raw English message rather than
blocking. A test demonstrates a customer with the type enabled receiving nothing.

recordOnlyEvents: a third routing class — stored and recorded, never mailed.
recovery_unit_capture_failed moves here: it is the record, the digest is the
notification. A register rather than downgrading severity to info, which would
relabel a genuine failure as informational everywhere it is queried.

cooldownRunSuffix: a sibling of cooldownTierSuffix, not a branch inside it, so
tier keeps byte-identical semantics and R-97a's tests are untouched. It makes
the cooldown effectively inert for the digest, which is the intent — a digest is
already rate-limited by construction; the refresh sweep sends no run_id and so
stays under the ordinary hourly cooldown.

The email renders as a list, not a JSON blob. An absent space reading renders as
unavailable, never as zeros.
2026-08-03 13:46:48 +02:00
admin 179dd79882 hub v0.89.0 — the two halves of decision D-c (R-167, R-158)
gates / gates (push) Successful in 7s
New OPERATOR-ONLY event type recovery_unit_capture_failed (controller
v0.191.0, R-158): in allowedEventTypes AND notify.operatorOnlyEvents.
Deliberately not a reuse of backup_failed, which carries customer copy and
sits in the controller's DefaultEnabledEvents — reusing it would email the
customer in Hungarian about a failure they cannot act on. R-158's own
proposal said backup_failed; D-c overrides it.

disk_warning/disk_critical lose their generic customerMessages entries.
Both were allowlisted, copy'd, default-enabled and checkbox'd with NO
producer anywhere; controller v0.191.0 becomes that producer and sends a
DYNAMIC Hungarian message naming the drive and its free space.
FormatCustomerEmail prefers the entry over the message, so keeping a static
entry would discard the label and the byte figures — the same reason
offbox_enlarge_blocked and disk_health_degraded have none. The deletion is
pinned by a test.

New notify.IsOperatorOnly so the api package can pin BOTH registers of a new
event type in ONE test; allowlisted-but-not-operator-only is invisible when
they are checked separately, and it is the defect v0.78.0 shipped. The
register itself stays unexported.

REUSE.md's "new event type" extension point rewritten: it told readers to
always add a customerMessages entry, which is wrong for operator-only types
and harmful for dynamic-message ones.

Tests 574 -> 579. Red-proof: removing the operatorOnlyEvents entry shows the
customer being emailed; the skipped/operator_only row is asserted as a
positive observable.
2026-08-02 23:19:13 +02:00
admin 0fc54e0122 hub v0.88.0 — the WAL that never was (R-172)
gates / gates (push) Successful in 7s
store.New opened the DB with `?_journal_mode=WAL&_busy_timeout=5000`, which is
mattn/go-sqlite3 syntax. The driver is modernc.org/sqlite, whose applyQueryParams
reads only _pragma/_time_format/_time_integer_format/_txlock/_inttotime and
IGNORES anything else WITHOUT AN ERROR. So the hub ran in rollback-journal mode
with busy_timeout=0 for its entire life while its own source said otherwise.

Surfaced as a false HOST STALE banner: in rollback-journal mode a reader excludes
a writer, so rendering an operator page blocks a host report; the hub 500s, the
agent waits its full 15-minute interval without retrying, and staleness fires at
30 minutes — two collisions is a false alarm plus an operator email. 13 collisions
in one pod lifetime; the alarm fired twice on 2026-08-02 for a host that was up
two days and reconciling throughout.

The observable that proved it: a 128 MB /data/hub.db with no -wal/-shm beside it
while the DB was open.

Fix: ?_pragma=journal_mode(WAL)&_pragma=busy_timeout(5000)&_txlock=immediate.
_txlock=immediate is not optional — database/sql's Begin() is DEFERRED, so a
read-then-write tx must upgrade its lock and a failed upgrade is
SQLITE_BUSY_SNAPSHOT, which busy_timeout does NOT retry; this store has 10+
db.Begin() sites and they are all write paths.

Every test asserts what the DATABASE reports, never the DSN string — a string
test would have passed for the whole life of the bug. Red-proof: restoring the
shipped DSN reproduces journal_mode="delete", the missing -wal, and the live
"database is locked (5) (SQLITE_BUSY)".

Operational consequence handled: a WAL DB cannot be copied by taking hub.db
alone — a bare `cat` opens cleanly and silently omits the newest writes. The
break-glass retrieval in operations/nodes.md used exactly that; it and the
recovery-inventory note are now WAL-aware.
2026-08-02 21:06:29 +02:00
admin 4cc123809c revert the Scenario B breakage — main is green again
gates / gates (push) Successful in 7s
The deliberate hostInstallVersion const is removed. It existed only to produce a real red
run (#3-#6) and the demonstrated alarm; R-94's deletion stands.
2026-08-02 16:26:30 +02:00
admin 3252d51104 SCENARIO B: deliberately break the hostinstall gate (reverted immediately)
gates / gates (push) Failing after 7s
Pushed with --no-verify ON PURPOSE: this simulates exactly the bypass that CI exists to
catch. The local pre-push hook would have refused this commit.
2026-08-02 16:21:35 +02:00
admin d319ae573e hub: delete the host-install version label (R-94) + invert hostinstall gate 1
The Setup tab said 'host-install 1.19.0' while the served script was 1.22.0, and had
been wrong since 2026-07-14. Deriving the number honestly is not possible: the Option-1
command downloads felhom-host-install.sh from the website at RUN TIME and the website
git-syncs main every 30s (R-110), so no build-time value in the hub can be true. R-94(a)
offered derive-or-delete; deleted, which removes the drift class instead of automating it.

- configs.go: hostInstallVersion const, pageData.ScriptVersion field and its assignment
  all removed; a NOTE in their place records why there is no constant here.
- customer_unified.html: the sentence now says the command always fetches the current
  installer, and renders no version.
- hostinstall_gates.py gate 1: the third assertion INVERTS — it used to require the hub
  const to equal SCRIPT_VERSION, it now asserts the hub carries no host-install version
  literal at all, matched in six code shapes across every .go/.html under hub/ (comments
  are deliberately not stripped: a // inside a URL literal would blind the scan).
- render_test.go: the assertion 'html contains hostInstallVersion' compared the constant
  to itself and passed at ANY value — demonstrated green with the const at 9.9.9 while the
  script was 1.22.0. Deleted, not replaced: there is no longer a version to assert.
- felhom-host-install.sh: COMMENT ONLY (SCRIPT_VERSION untouched) — it claimed the gate
  keeps the hub copy equal, an invariant that no longer exists.

Red-proofs: restoring the const fails the rewritten gate 1 (3 shapes hit); the old
render_test assertion passes at 9.9.9.
2026-08-02 15:16:01 +02:00
admin 670ec35ece hub v0.86.0 — Copy works without revealing, and every copy branch reports itself
Found by the operator, in the way that matters: it cost a real login.

The v0.84.0 Console access card shipped its Copy button DISABLED until a Reveal.
Clicking it did nothing, silently, so the clipboard kept whatever was already in
it — another host's console password from an earlier reveal. That got pasted into
demo-hp's PVE login, which failed with no explanation: the box logged a plain
`password check failed for user (root)`, the credential was never at fault, and
nothing on screen said the copy had not happened.

A copy button that silently no-ops is worse than no copy button. The operator
cannot tell "copied" from "did nothing", and the stale value left behind is a
VALID secret for a DIFFERENT machine — so the failure looks like a stale
credential and sends you diagnosing the wrong thing.

Copy now works without revealing, and that is the safer default rather than a
concession: the secret goes straight to the clipboard and never renders on
screen, so it cannot be shoulder-surfed or caught in a screenshot. Reveal remains
for when it must be read.

Three silent-failure branches closed, all in the same eight-line function:
  - not yet revealed        -> was a disabled no-op; now fetches and copies
  - navigator.clipboard absent -> was silently skipped; now shows it and says why
  - writeText() REJECTED    -> promise was ignored, so the operator believed it
                               copied; now shows it and reports the refusal

The success path names the host ("Copied demo-hp-bb76ea's root@pam password"),
because the clipboard is fleet-wide and every box has a different console
password — "copied" alone cannot say for WHICH box, which is the confusion that
produced the incident.

One retrieval path, shared: the endpoint is defined once (data-reveal-url) and
read back with getAttribute, so Copy cannot drift onto a different, unaudited URL
than Reveal. Server-side is unchanged — both buttons hit the same CSRF-gated
endpoint and both write the same recovery_credential_revealed event, which is
correct: the register records accesses, and a copy is an access.

Tests 566 -> 568. Red-proof: re-adding `disabled` reproduces the shipped bug.
2026-07-31 09:22:11 +02:00
admin e07d90f0f4 hub v0.85.0 — Network card: a host's addresses are visible at last
Pairs with agent v0.119.0 and is useless without it.

A managed box's LAN IP was not shown anywhere in the hub, because nothing
reported it — the host report carried no address of any kind. The only IP
reachable from the UI at all was the WireGuard one, on /offsite's peer table
keyed by pubkey, so an operator could go peer->host and never host->peer, which
is the direction anyone actually asks in.

The host page grows a Network card: every routable address the box holds, one row
per (interface, address), plus a WireGuard row. On demo-felhom that is vmbr0
192.168.0.162/24 and tailscale0 100.70.170.35/32 — with the PVE web console at
https://<the LAN address>:8006, the thing the operator wanted and could not get.

WireGuard is rendered as TWO facts, deliberately. WGAssignedIP is the hub's own
allocation (wg_peers, authoritative desired state); WGConfirmed is whether the box
reports actually holding it. Showing the allocation alone would make a peer that
was never applied look healthy — the same shape as reading a timestamp that
records an attempt as if it recorded a result.

The split is keyed on the ALLOCATION, not the interface name: wg-felhom is the
agent's current unit name, and a UI keyed on that string would silently
mis-render the day it changes.

An old agent renders UNKNOWN, never "no addresses". Below agent 0.119.0 the field
is absent from the wire, and an absent signal is not a negative result — the page
says so and names the version needed. Rendering an empty list there would have
stated something false about the host.

No new store table and no new ingest path: the report is already stored opaquely
and GetWGPeerForHost already existed with no UI consumer. This is parse + render.

The report fixture in the tests is the REAL wire — the addresses block copied out
of `felhom-agent --selftest=hub` on demo-felhom running 0.119.0.

Tests 559 -> 566; four red-proofs (inert view-model, unconditional confirmation,
the old-agent branch, and the drift case) each run, observed failing, reverted.
2026-07-31 08:49:44 +02:00
admin 1956e5d390 hub v0.84.0 — break-glass console credential on the host page
The credential existed and was not reachable when it was wanted. Every box has
had a strong random root@pam password since TASK G1, vaulted in the hub at day 0
and used for real during the sshd incident — but the only way to read it back was
a hand-written curl carrying the global operator key, a secret kept out-of-band.
In practice the PVE web console on a demo box felt locked.

The host page grows a Console access card: presence + username + set_at by
default, Reveal fetches the plaintext on demand for 60 s with a Copy button.
Masking clears the JS variable, and also fires on a second click and on
visibilitychange. A host with nothing vaulted says so, and says why.

The secret is NEVER rendered into the page, and that constraint shapes the
change. The render path uses a new store.GetHostRecoveryMeta whose struct and
SELECT both omit the secret column, so it is structurally incapable of carrying
one. The plaintext crosses the wire only in the response to POST
/hosts/{id}/reveal-recovery-credential (Cache-Control: no-store, CSRF-gated at
the ServeHTTP level; POST precisely so that gate applies and so no secret is
retrievable by URL alone). Deliberately NOT the customer page's data-secret
widget, which embeds the plaintext on every load.

A delivered reveal writes one recovery_credential_revealed event on the host's
customer timeline (info, source hub, Hungarian) via SaveEvent alone — no
dispatcher, nobody emailed, the log_tail_requested shape. Two reveals write two
events: the register records accesses, not states. A 404 is not an access. An
unbound host reveals fine and writes no event; the [INFO] hub line, carrying the
username and a length only, is then the record.

The global-key API path is untouched by design — it is the route for when the
hub UI itself is broken, and coupling it to the session layer would delete the
independence that makes it a fallback.

Recorded as a real trade: the hub session password alone now unlocks console root
fleet-wide, where retrieval previously also needed the global key. Accepted for a
single-operator, HU-geo-fenced hub that already stores these passwords in
plaintext at rest (CONTEXT.md ruling S-4). The plaintext-at-rest half is filed as
R-133 — every hub DB backup is a fleet-wide console-credential dump.

Tests 550 -> 559; four red-proofs (page leak, audit event, CSRF gate, route
order) each run, observed failing, and reverted. The route-order proof is a seam
test driving ServeHTTP: a handler-level test cannot see that defect, because the
handler is correct and simply never runs.
2026-07-31 08:19:36 +02:00
admin acfc2b7e95 R-109 + R-122: the recipe assembly stops dropping sections (hub v0.83.0)
AssembleDRRecipe's hostHalfShape/appHalfShape are ALLOW-LISTS, not the
forward-compat their comment advertised: a section an emitter adds is silently
discarded until it is named in both the shape struct and AssembledRecipe. No
error, no log, no failing test.

R-122 (found this session): that already happened and shipped. The controller
has emitted offsite_restic since fork-4 — the offsite recovery LOCATION — the
hub stored it for all three real customers, and appHalfShape never listed the
key, so no delivered recipe has ever contained it. It stayed green because the
fixture drAppHalf is hand-written and omits the field.

R-109: the agent's new backup_target is a new top-level host-half section and
would have been dropped identically, making the fix read as shipped while
changing nothing an operator can see.

3 tests built on halves read verbatim out of the live dr_recipe table, plus
2 red-proofs (each mutation asserted to have landed). vet rc=0, suite rc=0, 17 ok.

Registers: R-106 + R-109 dispositioned; R-105/R-106 were READY in ROADMAP with
no OPEN-ITEMS row (→ R-123, registered); R-124 filed on the "root" spelling.
2026-07-30 13:13:56 +02:00