adb8d904ca1a1c679c0eb64ca22dc9e76cb969a6
249 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ccd915ff34 |
hub v0.125.0: the customer delete lists the Cloudflare items to remove by hand instead of promising it (R-688)
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
2d24931597 |
hub v0.124.0: a thin pool is critical at 90% (data or metadata), one alarm per pool per 6 h (R-672)
gates / gates (push) Successful in 24s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
99c0709cbe |
hub v0.123.0: app_stopped_unhealthy (decision 28); 09 decisions 26-28 recorded
gates / gates (push) Successful in 31s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
e4d45a8f72 |
hub v0.122.0: app_hold_no_whole_copy — allow-listed, operator-only, per-app cooldown (R-659)
gates / gates (push) Successful in 28s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
511969928d |
hub v0.121.0: app_oom_storm allow-listed, operator-only, per-app cooldown (R-636)
gates / gates (push) Successful in 24s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
d3b284863e |
hub v0.120.0: app_update_undone/held reach the household, per app, in its language
gates / gates (push) Successful in 26s
Allowlisted, not operator-only, seeded for new households and added once (add-only) to every existing enabled_events row. mail.event entries in hu and en name the app from details.stack_name. Per-app cooldown on both the operator and the household leg. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
e02bc03819 |
hub v0.119.0 — English households get English words for their codes (R-597); R-596/R-598 closed
gates / gates (push) Successful in 24s
The setup code and the owner passphrase now follow the household's language,
one word longer in English so the entropy never drops (setup 3 hu / 4 en,
passphrase 5 hu / 6 en). List and count are chosen together so a caller cannot
pair an English list with a Hungarian count. Hungarian is byte-unchanged.
Three claims in the row were wrong and are recorded as such:
- the RECOVERY CODE is minted by felhom-agent from the EFF list and has
always been English; the hub does not own it and no row was added.
- no claim mail states a word count; the only count wording was the bind
page's passphrase hint, whose English half is now count-free.
- the proposed phone-safe filter removes 68% of the list (5270 of 7772
words) and was measured, then declined, with the reason in source.
Also: guide_quote_gate binds the English volunteer guide's three quoted
messages to the controller's English bundle — nothing did, so the guide would
have gone on quoting Hungarian after the fix. Seven decoys, all convicting,
including the name-for-fact one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
|
||
|
|
a2c52ebf2a |
hub v0.118.1: the test mail follows the language too (R-558)
gates / gates (push) Successful in 23s
Found during v0.118.0's own live proof, which is the only reason it was found: I went to press "send test notification" for the English demo box and read sendTestEmail first. It had its own hardcoded Hungarian subject and body and never went through FormatCustomerEmail, so it was the one customer mail v0.118.0 did not localise — and it is the only customer mail an operator can trigger on demand, which makes it the one most likely to be used to check whether the localisation works. Pressing the button for an English household would have answered that question wrongly, and convincingly. The two sentences are extracted byte-for-byte into the bundle, so the Hungarian test mail is unchanged. Red-proofed against the hardcoded version. The general form worth keeping: the surface you would use to CHECK a feature is the one most worth checking first. A broken instrument that reports success is worse than a broken feature. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
9167cf53af |
hub v0.118.0: the household's e-mails follow the household's language (R-558 Part A)
gates / gates (push) Successful in 23s
The hub has written every customer e-mail in Hungarian whatever the box was set to. The box has published its language since controller v0.247.0; nothing read it. Now it does. Nothing an operator reads changes. The Hungarian mails are byte-identical, and that is a diff rather than a reading: 56 goldens per language captured from v0.117.0 BEFORE any string moved, and all 56 Hungarian ones pass unchanged after every sentence was routed through the new bundle. - internal/i18n: flat bundle, 79 keys, hu authoritative + hu fallback, ceiling 0. - customerMessages/severityLabels are DERIVED from the bundle, so a sentence is written in one place and all 40+ tests that read those maps still work. - Language order: last reported -> created-with -> hu. reports.language defaults to EMPTY, never hu: "never told us" is not "chose Hungarian". - message_customer on POST /api/v1/event, additive and optional forever, for the sentences the box composes and the hub cannot translate. - The bind page is per-language, and its `expired` state stays Hungarian: it is the state an unknown token lands in, so rendering a real English customer's token in English would make the LANGUAGE answer what the TEXT refuses to. Two defects found inside the release: - R-581: the newest report was picked by received_at, which has SECOND granularity, so same-second reports tied and the winner was arbitrary. Ordered by the autoincrement id now. GetCustomers() still has the shape - row open. - R-582: the English copy-guard stems, ported word for word from Hungarian, convicted 141 honest sentences. The English claim is a phrase with a modal. R-555 closed: the language allowlist entry is out of wire_contract_gate.py. hub_copy_gate.py follows the sentences into the bundle - without that it would have scanned four files that no longer hold any customer text and reported success. Three new decoys incl. an innocent control. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
37ae31fd44 |
hub v0.117.0: the status follows the configured threshold; slow crash loop and interrupted restore events
gates / gates (push) Successful in 21s
R-549 (operator ruling A): controllerStatus hardcoded 30m/1h while both staleness checkers and hostStatus read alerting.stale_threshold. Moving the threshold to 45m would have painted a customer amber 15 minutes before the alarm could fire - the second definition rollup.go's header forbids. It now reads the same value, down at 2x. Both 'checker initialized' log lines print the threshold, which no line did before. R-539 (ruling 3 of 2026-09-16): controller_slow_crashloop (warning, operator-only), minted when the agent's slow_crashloop_since moves, with the fast sibling's first-sight rule. R-550: restore_interrupted (warning, for the household) allowlisted with a Hungarian customer message. Red-proofs, each seen failing then passing: the status test with the old hardcoded numbers; the checker test with the movement branch removed; the operator-only test with the registration removed; the household-message test with the Hungarian entry removed (asserted on the SUBJECT - the body legitimately repeats the raw message, which my first version of the test mistook for a fallback). go build/vet/test ./... green, 18 packages. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
3738dfc548 |
hub v0.116.0: every new customer starts WITH the off-site copy (operator ruling)
gates / gates (push) Successful in 19s
Off-site is ON by default for a new customer — shared, 100 GB soft quota prefilled, the checkbox kept so an operator can opt a customer out. The reason is this repo's own [FACT]: the whole-guest tiers do not carry the data drive and a Tier-1 unit has no file leg, so with this unticked a one-drive box keeps NO copy of the household's own files. Measured on a fresh box the same day. The quota is prefilled because the fill warning only fires when quota_gb > 0. Also registers controller v0.244.0's app_deploy_started / app_deploy_failed in both allowedEventTypes and customerMessages, per the rule that the two move together. Red-proofed: dropping the default fails the new render test. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
926723749d |
hub v0.115.0: host_* mails skip the quiet hour (ruling 2, R-529); ruling 1 recorded (CC may sign agent_update until the first paying customer)
gates / gates (push) Successful in 23s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
07959e61b5 |
hub v0.114.0: self-bind auto-send while a customer waits for a box (R-509); node_* bypass the quiet hour (ruling 2026-09-15); PBS re-issue adopts an endpoint token (R-511); controller supervisor events (R-523); event registers
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
6fd8c87516 |
doorstep: console is Felhom's (ISO 1.27.0 source), passphrase hand-over copy (hub 0.113.0 source), rulings
gates / gates (push) Successful in 19s
Phase 0: the public ISO never auto-installs by construction (no answer.toml, G1); the operator re-affirmed the interactive installer 2026-09-14. - felhom-bootstrap.sh: mask pvebanner.service, write a Hungarian /etc/issue (no :8006 admin URL); pairing banner names the Tulajdonosi jelmondat and paints through the CONSOLE_DEV seam (R-496). Harness: 8 checks, red first; fake hub now sends a pairing code (the banner was never tested, R-502). - hub: created flash + Credentials block tell the operator to hand the phrase over; the self-bind mail names the operator (R-497). Tests red first. - iso-release-gate G14-G16; domain ruling in 01-topology + CONTEXT; R-494 narrowed to P3; R-502..R-504 filed; volunteer guide and day-0 A.2 aligned. ISO_VERSION 1.27.0 (not built, not published). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
f181efd6a7 |
hub v0.112.0: a floor carries a declared MinAgent past the golden (R-472)
gates / gates (push) Successful in 18s
Operator ruling 2026-09-13. Above the vouched golden, a floor saved with a declared MinAgent is served under the same agent comparison; an undeclared one is still held beyond the golden. The declaration is stored beside each floor as FLOOR=MINAGENT so it never carries to a later floor. Both floor forms require min_agent above the golden (flash floor_needs_min_agent, nothing stored). The Hosts page and the API log name the source. Vouch path and R-120 gate untouched. Scenarios A-E tested; red-proofs A and C in documentation/audits/rulings-r472-r475-2026-09-13/. |
||
|
|
db38f4c800 |
hub v0.111.1: the alarm stops promising a rescue that does not exist, and the arc is closed for beta
gates / gates (push) Successful in 17s
R-434 CLOSED — and the row's own "blocked on R-433" verdict was wrong, which is the point.
The fix is a DELETION, not a replacement: withdraw the promise instead of swapping it for a
new one, and the sentence is true under every possible answer to the provider questions, so
it never needs a second rewrite. A replacement would have been blocked; a withdrawal is not.
was: "...still hold the older copy, so this is recoverable file-by-file; it is NOT
confirmed data loss. Check whether a deletion ran on the box before restoring."
now: "...still hold the older copy. The route back out of them is not yet established,
so treat this as neither confirmed data loss nor confirmed recovery. Get in touch
before restoring anything, and check whether a deletion ran on the box."
It must not swing the other way either: "your backups are gone" is still usually false.
Clause (a) — the box cannot WRITE into the snapshot area — stands and is re-confirmed.
Tests: offsite_r434_test.go, three, all driving the production path so they assert the
sentence an operator RECEIVES. ASCII-only fragments, positive and negative controls.
RED-PROOF: restoring the v0.111.0 sentence failed all three, on every fragment, with the
offending sentence printed. TestR431_FiresOnAMassDeletion asserted "NOT confirmed data
loss" and caught this fix correctly; its wording fragment is REMOVED rather than updated,
so the wording keeps ONE home.
R-435 written into the detector's own documentation, no threshold changed: it sees a mass
deletion, not one app being wiped (69 across 9 apps -> ~35 needed, one tag is ~9, and
forget --prune groups by host,tags). Says explicitly not to lower the numbers.
THE STOPPING LINE, in all three places — register, 07 section 8 head, STATUS.md.
Deferred set ENUMERATED, not described: 07 section 8 rows 4, 8, 9, 10, 11 (+11b), 12,
each tagged [BETA-DEFERRED]. A number in the brief was wrong and is corrected in place:
six rows are DEFERRED, ELEVEN carry a blank RTO (4,5,8,9,10,11,11b,12,13,14,15); the other
five are blank for reasons that are not deferred work, and row 15 is an open DEFECT (R-104)
that the stopping line does NOT cover. NO STATUS MOVED — nothing was proven today.
Two provider questions drafted, not sent, no API called (11-D stands):
documentation/runbooks/provider-questions-2026-09-01.md, linked from R-95 and R-433, and
tracked by a dated DUE-CHECKS row (2026-09-15) — the 2026-07-27 check that sat unconfirmed
for 36 days is the scar that block exists for.
R-95, R-433 BLOCKED-ON-PROVIDER. R-95's one-day demotion on a clause that did not hold is
recorded; the proposal to rank it back near the top is stated and NOT acted on. R-430 marked
LATENT with its trigger: it becomes live the moment delete is withdrawn, so it is a
precondition on the R-95 build, not a follow-up. The stale ranking paragraph ("armed",
"zero snapshots") is corrected in place, order unchanged.
Register 621 -> 688 lines; 181 rows throughout; open-state 170 -> 169.
No controller or agent change. No golden owed, no floor change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
|
||
|
|
30681764cb |
hub v0.111.0: notice a deletion within a day (R-431); correct R-429; re-scope R-95
gates / gates (push) Successful in 17s
THE RECORD WAS TELLING A WORSE STORY THAN THE TRUTH FOR TWO MONTHS, and my own probe is why.
R-429 CORRECTED. Yesterday's spike searched for a directory called `.snapshots`. The vendor documents
the path as /.zfs/snapshot. The probe's CONTROLS were sound and its SUBJECT was wrong, so "not found"
was true and meant nothing. Re-probed at the documented path on both boxes, with controls:
- /.zfs lists (shares, snapshot) from inside the jail;
- a write into /.zfs/snapshot is REFUSED - dest open ...: Failure - while the identical write to
the account home SUCCEEDS and was cleaned up.
That is the append-only property PROVEN rather than cited, and it is the sentence the whole re-scope
rests on. Seven daily snapshots are confirmed in the panel. The mitigation works. What remains true,
and was always the actual finding: the row claiming it had no R-number, its "confirm tomorrow" went
36 days unanswered, and the DUE-CHECKS block built for that class was empty. THE FINDING WAS NEVER THE
SNAPSHOTS - IT WAS THAT NOBODY COULD TELL.
R-95 RE-SCOPED: the box can delete its LIVE repository but cannot write to the daily snapshots of it,
so a deletion costs at most one day plus a per-file recovery - not open-ended loss. The ranking is
Viktor's; it has been #1 since July on the old story.
R-432 FILED: a sub-account sees /.zfs/snapshot EMPTY while the same box holds seven snapshots, so
per-file recovery is operator-only today. One panel read settles whether a NAMED snapshot can still be
entered, which would make it product-reachable.
R-431 SHIPPED. Third signal in OffsiteChecker. On the hub deliberately: a detector on the box is one
the deletion can silence. Threshold REASONED, not invented - over 12 898 reports every decrease lands
on ZERO and predates stats_known, and in the stats_known window there are none, so observed churn gave
nothing to calibrate against. Retention cannot halve a total; a mass deletion goes to ~0. Hence: more
than half, and at least 5. Guarded by StatsKnown (R-331), the declared State (R-204) and run success
(R-100 - whose lesson lives in this very file).
ACCEPTANCE: 9 009 real report points replayed through the detector produced ZERO alarms.
Three red-proofs run. The escalation one only became real after the first version was found HOLLOW -
it re-swept the same report, so the baseline had already moved and the latch was never consulted.
07 row 10's status is NOT moved: the write-refusal is measured, but the recovery ROUTE has never been
walked, which is what PARTIAL means.
|
||
|
|
1aeaa30c28 |
hub v0.110.0: allowlist + operator-only for offsite_proof_empty (R-87)
gates / gates (push) Successful in 15s
Controller v0.231.0 adds a nightly job that proves an app's newest off-site snapshot still CONTAINS that app's data. When it finds one that does not, it emits offsite_proof_empty. Two register lines, both load-bearing and both in this commit: - allowedEventTypes - an unallowlisted type is answered 400 and VANISHES, so this entry is what makes the alarm exist at all. - operatorOnlyEvents - a missing customerMessages entry is NOT a routing block (FormatCustomerEmail falls back to the raw message, the v0.78.0 defect that register was built for). A customer can take no action on a hollow recovery unit. DELIBERATELY NOT reusing backup_integrity_failed, which is the nearest existing type: it means THE STORE IS DAMAGED and carries the Hungarian template saying so. Here the store is sound and the CONTENT is absent - different cause, different action, and telling a customer their backups are damaged when they are not is the more expensive mistake. Same asymmetry looksLikeRepositoryDamage is shaped around. DELIBERATELY no customerMessages entry (the controller's dynamic Hungarian names the app and what is missing; a template would discard it) and DELIBERATELY not in perAppCooldownEvents (a fenced act - this job proves ONE app per night, so the coarse hourly cooldown is already the right grain). This widens the R-87 task's stated repo scope to felhom.eu/hub/. The reason is recorded in felhom-controller/CONTEXT.md ruling 4 rather than left as an unexplained diff. Hub green gate: go build/vet/test all pass, 18 packages. |
||
|
|
36f8630020 |
R-341 check taken, and the golden-currency bypass declared
gates / gates (push) Failing after 16s
Two gates blocked the R-331 hub push. One is FIXED, one is BYPASSED, and the difference is stated rather than blurred. FIXED -- due-checks (R-341, 5 days overdue). The +7d measurement was TAKEN on ep0 rather than deferred again. Precondition passed: proxy still MainPID 551655, ps -o lstart= still 2026-08-18 09:51:04, NRestarts=0, so this is the same proxy generation as t0 (anchor is ps, not ActiveEnterTimestamp, which reads 03:54:54Z here -- R-346's trap). Result: fd = 17. Not 17 more -- seventeen TOTAL, exactly the documented baseline, against 405 at the first check. Socket histogram: one LISTEN, ESTAB 0, CLOSE-WAIT 0. The verdict is UNANSWERABLE, not "the upgrade fixed it". R-341 asks whether the PBS 4.2.5-1 upgrade changed the fd slope; inside this interval we removed the leak OURSELVES (R-344, agent 0.130.0, now live on both boxes). A slope of ~0 measures our fix, not the upgrade, and reading it the other way would credit a changelog that was read in advance and found to contain no such mechanism. The perturbation pre-registered for this window was Phase C at ~3%; the actual perturbation was the removal of the entire phenomenon. Row closed as moot. What it DOES establish is worth more than the original question: twelve days after the R-344 fix, same proxy generation, no restart to hide behind, ep0 sits at baseline with zero established connections. R-336's ~323-day runway concern retires with it. BYPASSED -- golden-currency. Controller v0.224.0 and v0.225.0 are released and the newest golden bake carries 0.223.0, so a machine installed right now gets neither. The gate is RIGHT. This push therefore uses `git push --no-verify`, declared here, in hub/CHANGELOG.md, in REPORT.md and on R-242. A BYPASS, not a waiver: the gate offers a waiver only for a release that DELIBERATELY needs no golden, and these need one. The operator was asked and ruled bypass-now-bake-later, on the ground that neither fix bites a day-0 box -- R-330 is a nightly false alarm about apps a new box has not installed yet, R-331 is a hub display over backups a new box has not taken yet -- and both arrive by self-update. That ground is recorded because it is what to re-check: it does NOT extend to a release changing first-boot behaviour. OWED: bake a golden carrying 0.225.0 and vouch it (RUNBOOK-manual-build.md 4.1, three-field change, MinAgent 0.129.0). Fourth bypass of this gate, and the gap is now two releases wide rather than one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM |
||
|
|
f5c9411e5e |
R-331 (hub half): the Backup card reads offsite, not the dead backup fields (v0.109.0)
The customer page's Backup card read `Snapshots 0 / Repo Size 0 MB / Integrity
Unknown` for EVERY customer, indefinitely. Measured on demo-hp 2026-08-30 while
that night's controller log said `[offbox] backup OK: 8 app(s) backed up, 67
snapshot(s), 2m14s` and the box held snapshot_count:67, repo_size_bytes:
140829678, stats_known:true.
A card reading "no backups" over a working backup is worse than no card -- the
R-88 direction of failure (degrade to NO BACKUP rather than to UNKNOWN) on the
one screen that answers "is this customer protected?".
The data was never missing. The card rendered the report's `backup` object,
whose snapshot/size/integrity fields have had no producer since slice 8C. The
live numbers are in the `offsite` object, which THIS PACKAGE already reads for
the Offsite page and which monitor.OffsiteChecker already alarms from. Proof the
bytes were arriving: the Offsite page rendered demo-hp's usage as 0.1 GB from
that very object while the Backup card said 0 MB. So this is a render fix over
an existing feed, not a new pipeline.
Not a one-line swap, because snapshot_count:0 means two opposite things --
"holds nothing" and "never measured". R-225 measured that confusion one layer
down. backup_card.go resolves a three-way ruling in Go (a {{if}} chain over
map[string]interface{} float64s cannot keep the absent/zero distinction the card
is entirely about):
no offsite object -> "No off-site data reported", and says explicitly that
this is NOT the same as "no backups"
disabled + state -> names the blocker (needs_credential)
stats_known:false -> em-dash + "never been measured". NEVER 0
stats_known:true -> the real numbers, INCLUDING a real 0
A pre-v0.225.0 controller sends no stats_known -> false -> "unknown". That
direction is pinned: upgrading the hub ahead of the fleet must not report every
un-upgraded customer as having zero backups.
The Integrity row is DELETED, not re-sourced: nothing produces it, the
controller runs no integrity check, and NotifyIntegrityOK/Failed are called from
nowhere.
RED-PROOF: restore the pre-fix card markup -> all four tests fail, reporting 67
and 134.3 MB absent from the rendered page and the Integrity row present. The
tests drive handleCustomerUnified and grep the HTML on purpose: the defect was
the template's choice of source object, so a test one layer below it would have
been green against the shipped bug.
Green gate clean: 18 packages, rc 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
|
||
|
|
45659bdc5a |
hub v0.108.0: deploy the per-app cooldown grain (R-389)
gates / gates (push) Successful in 17s
Image built and pushed to the registry BEFORE this manifest bump lands, so a sync can never point at a missing tag. Auto-sync is off; the sync that follows is deliberate. Never kubectl set image. |
||
|
|
2fc4a15fa3 |
R-389: key the operator cooldown per app for app_start_failed; gate 11 makes an unfiled observation refuse the push
gates / gates (push) Successful in 16s
The cooldown key was customerID:eventType plus the tier and run suffixes, and none of them names an app, so every app going down inside the same hour collapsed onto one key and only the first was mailed. Measured on demo-hp: bookstack sent 09:27:51, privatebin suppressed 09:31:51 under key=demo-hp:app_start_failed. cooldownStackSuffix is the third sibling of cooldownTierSuffix and cooldownRunSuffix, and separate for the reason the second one's docstring already gives: the existing two keep byte-identical semantics for every type that uses them. It is ALLOW-LISTED to app_start_failed and takes the event type as well as the details, unlike its siblings, and that asymmetry is the safety property. The backup family's cooldown is coarse ON PURPOSE (R-97a, R-182) so one full disk sends one digest rather than one mail per app - and crossdrive_failed is severity error, reaches the operator leg, and carries stack_name through a DIFFERENT struct, so a payload-shape rule would have split it silently. The hour itself does not change. Gate 11 refuses a push whose REPORT.md carries an observation with neither `FILED: R-NNN` nor `NOT-A-FINDING: <reason>`. It deliberately does NOT accept a passing mention of some other R-number: the lost item cited R-182 as an analogy, so "cites a register row" would have passed the very item the gate exists to catch. That discrepancy with the spec is recorded in the gate's docstring. Registered here and in the controller and agent runners. NOT in the catalog runner - it has no shared-gate mechanism and appends --all to every gate; filed as R-391 rather than left as a sentence, which is this session's lesson. PROMPT-TEMPLATE.md §15.9 corrected: "documented, NOT acted on" was the wording that invited the gap, and it now names the markers and points at the gate. R-390 filed for the golden-bake runbook's missing `pveam update`. Hub tests 709 -> 716. |
||
|
|
68a9f5475c |
hub v0.107.0: the hub rewrote a severity and said nothing (R-387); golden 0.223.0
gates / gates (push) Successful in 17s
One handler, two fields, opposite discipline. An unknown event_type is rejected with a loud 400. An unknown severity was rewritten to "info" without a word - and severityNotifies drops "info" before BOTH legs, so the event was stored, answered 200, and mailed to nobody. Two shipped features went out that way: DiskAlertKind.Severity emitted "warn" until controller v0.215.0, app_start_failed until v0.223.0. Measured on the live hub DB today: 91 app_start_failed events stored all-time, ZERO notification_log rows before this session - not one, on any channel. The mechanism built to catch this class was structurally blind to it: the dispatcher's `unrecognized severity` line cannot execute for anything arriving over the API, because the coercion one line earlier guarantees the value it looks for cannot arrive. The coercion STAYS - a rejected event is a lost event, and losing an alarm is worse than mis-routing one. Only the silence is fixed: a WARN naming the customer, the event type and the rejected value. The dispatcher branch is KEPT, not deleted as dead, and the reason is evidence rather than caution: cmd/hub/main.go wires dispatcher.ProcessEvent DIRECTLY as the monitor.EventNotifyFunc for the staleness, host-staleness and offsite-box checkers, which never pass through the handler. For those it is the only severity guard there is. All 90 severity literals in internal/monitor are already valid, so the guard is silent because the producers are correct. Test count 702 -> 709. Red-proof seen failing: delete the WARN line and the coercion test fails with "the hub rewrote a severity and said nothing". Golden 0.223.0 baked and published (sha 9eaf39ac3921...), round-trip HTTP 206. Vouching is the operator's act and was not done here. |
||
|
|
ab2262c91c |
hub v0.106.0: report loss of visibility into the off-site stores (R-339)
gates / gates (push) Successful in 14s
THE GAP, measured not supposed. On 2026-08-18 ep0's PBS proxy was wedged for
9 h 37 m and the hub emitted NOTHING on the operator channel. Both box
checkers hold their last snapshot and return silently on a failed fetch --
correct for a FILL signal, since a missing reading must never be read as 0%,
but it makes a dead off-site endpoint and a healthy one indistinguishable.
The only mails that morning came from the boxes' own backup failures, and
only because the WEEKLY offsite run happened to land inside the window. Two
days earlier nothing would have fired at all.
REACHABILITY is now a second, independent signal on both checkers:
consecutive failed fetch windows, reported past a default 3 windows
(~30-45 min) as pbsdr_box_unreachable / offsite_box_unreachable (warning) on
the customer-less pbsdr-box / pool-box scopes, each with a paired *_recovered
all-clear. Tunable via alerting.box_unreachable_windows (0/invalid -> 3).
THE FILL LOGIC IS UNTOUCHED. No threshold, throttle, band or escalate-once
behaviour changed; a degraded read still drives no transition.
Three decisions a later reader would otherwise "fix" back, so each is
argued in-code:
- the unreachable event REPEATS rather than escalating once. The band shape
would give exactly ONE mail at ~minute 30 of a nine-hour outage, and one
mail is missable. It leans on the dispatcher's 1 h operator cooldown to
become an hourly "still blind" heartbeat.
- ErrUsageUnsupported is NOT blindness: an old ep0 answers "no such op",
which means we reached it. Counting it would alert for days on a healthy
pre-update endpoint.
- born-blind is reported: the counter is not gated on having a snapshot, so
a hub restarted INTO an outage still speaks. last_ok is OMITTED rather
than zero-valued -- a fabricated timestamp reads as "it was fine until
then".
Both recoveries are severity "info" and severityNotifies drops "info", so
they are registered in recoveredPairedDownTypes or the operator hears that
the tier broke and never that it healed. A cross-package test drives
ProcessEvent and asserts an actual operator MAIL, not a map entry -- a green
checker test proves nothing about the seam (agent v0.91.0 shipped fully green
with SetAuthSink never called).
Tests: box_reachability_test.go (Scenarios A-F) + dispatcher_box_reachability
_test.go (wiring). Three red-proofs run and reverted, each seen failing with a
message naming the right cause: threshold 3->1, the sentinel counter guard,
the pairing entry.
Register: R-339 filed and marked SHIPPED (PROVEN-LIVE still owed -- no real or
constructed outage has exercised the emit path, and one cannot be manufactured
against Tier-2 ep0). R-340 filed: the reachability read rides ep0's LOCAL API
daemon, which the incident explicitly cleared, so this check would have shown
GREEN for all 9 h 37 m -- the honest boundary, recorded rather than glossed.
R-336's next-step corrected: pvestatd's interval is NOT tunable (Proxmox staff
have said so); the only lever is disabling the storage entry, which collides
with the agent's consume-the-one-time-secret path. Doc-only, no agent code
touched.
|
||
|
|
b03a105375 |
hub v0.105.0: the third name, a machine told to be quiet, and a guard for the hub's own words
gates / gates (push) Successful in 17s
Hub only. No controller change, no agent change, no wire change — nothing to bake. demo-hp untouched: the operator is re-deploying it this evening. R-323 — the five-word phrase is „Tulajdonosi jelmondat". It was „Visszaállító jelszó": one word from the name retired last week, and false besides — it restores nothing, it proves the account owns the box being bound. Five sites, all in the hub; felhom-controller and felhom-agent carry the name nowhere, so no halt and no bake. Both suggested names were rejected with reasons: „Fiókjelszó" would collide with the dashboard login (a DIFFERENT real secret), and „Összekötési jelszó" would leave the two factors on this page separated only by kód-versus-jelszó — the exact shape being removed, since the other factor is the „Párosító kód". The chosen name differs on both axes, stem and noun. Naming only; the acceptance pin drives the real handler. R-324 — the hub's customer copy is under a guard for the first time. Retired names banned across all 95 hub files; retrieval stems registered in four declared customer surfaces. The selftest found a defect in its own instrument on the first run. One shared vocabulary in scripts/, drift-checked into the controller gate rather than copied (R-325 removes the scaffold). R-321 — a machine we told to be quiet is no longer reported as dead, and it was two doors, not one: because the state is RECORDED rather than deleted, the morning deadline check can skip it too. A deleted state returns "", which is not "down" — R-195's shape returning through a second door. The clock runs from the report the hub can see, so re-enabling starts it there and emits no recovery for an outage that never happened. Three red-proofs; the one that matters showed a genuinely dead machine sitting at "disabled" when the suppression was made unconditional. R-326 — "which claims are unproven" is answerable by a command now. The nine I have been repeating was the count of claims the 9 August pass DOWNGRADED, not the count of unproven ones. The real figures: 55 claims, 23 walked, 32 not — and only 6 of those 32 cite evidence. Its first run found a stale claim (R-327). |
||
|
|
4d6ec7c7bb |
hub v0.104.0: the guest network gets a reader (R-319), and the hub half of the naming (R-295)
gates / gates (push) Successful in 14s
Four paper debts and one fact given a reader. Hub-only — nothing to bake. A4 — the entry about "the tester's machine" named a risk correctly and labelled it in a way that invited deleting it. Established from the hub's own store: `peti-felhom` is a REAL machine (482 reports, 2026-02-27 → 2026-07-15, a named person's own box) and the 3.6 GB with no key and no backup is real. `david` → `tester-1` is a DIFFERENT record with no host, no escrow and no report, ever — deleted 07:55:49 and re-created 07:56:47 this morning. The prompt's premise conflated the two; the register now says which is which. A1 — R-312/R-313/R-303 recorded as DECIDED with their re-open triggers, and moved out of STATUS's "Waiting on you", which is now empty. A3 — day0-install §C.1 said pushing the installer publishes it. It has not since R-110. Corrected, with the two manifest pins named and an outside-verification command; the one copy that repeated it (a dated audit, true when written) carries a superseded note. A5 — standing rule 5: evidence comes off the machine at the end of the phase that produced it, before any revert. Earned twice in three days on the same box at the same point (R-320). Four homes, plus what to do when it is already gone. R-295 hub half — „Beállító kód" everywhere; „Visszaállító kód" retired. New `reenroll` mail kind so the mail names the page a REBUILT box actually shows („A szerver beállítása"), not the „Elfelejtett jelszó" page it has no login screen to reach. Naming only; the acceptance pin proves the secret is untouched. R-319 — the hub models `guest_net` after 23 days of receiving and discarding it. The signal is `heals_last_hour`, not `state`: a guest the watchdog keeps repairing reads healthy between repairs. `heal_succeeded` decoded too (R-260's lesson). Unknown is never drawn as healthy — three absences, three sentences. No alarm, deliberately. Three red-proofs, mutations asserted applied. Wire-gate checked tags 182 → 190. B1 — the operator's 2026-08-12 dispositions were NOT in the register; they are now. Third allowlist kind for the five ruled "no reader wanted"; `reporting_disabled` reclassified redundant. 8 read · 5 deliberately unread · 1 redundant · 6 still owed. Also filed: R-321 (a deliberately-silent box still alarms stale/down — the checker is age-only, and decoding the flag would not have fixed it), R-322 (the claim guard has never scanned the hub; a hand scan returns zero, so it is a scope gap, not a defect). |
||
|
|
6362bb6cb6 |
hub v0.103.0 — a host can read the packages we kept for it (R-311)
ListSupersededEscrow had zero production callers for nineteen days. It is the only reader of a retained identity_blob, so the retention shipped in v0.93.0 was material the product could not reach - proven on the fixture 2026-08-12, where a code that opens a retained package was answered as a code that opened nothing. New GET /api/v1/hosts/<id>/escrow/retained: self-scoped exactly as the current-row GET, same recovery-mode gate, same audit event written BEFORE the bytes leave, capped at 16. Rows with a NULL identity_blob are WITHHELD and returned as unopenable_count. They retain the PBS key, not the repository password, so they can never open what the caller is asking about; serving them would have the agent try packages that cannot succeed and would let the screen claim an earlier package is openable on exactly the boxes the original defect hurt. The count is returned because their existence is load-bearing and underivable. The trade, stated rather than waved through: the hub still cannot read any of it - sealed bytes in, sealed bytes out, no decrypt path, no recovery code ever held. What widens is volume, bounded by self-scope, the recovery-mode gate and the cap. The response is a NAMED TYPE, not a map, so the wire-contract gate can resolve it; the wire is declared as a fourth ROOT and the gate now checks 182 tags rather than 174. A positive control shows that check is name-presence, not decodability - filed as R-315 rather than reported as coverage. Six tests through the real endpoint; four red-proofs asserted applied. |
||
|
|
b55fc17d82 |
hub v0.102.0 — refuse to vouch a version that cannot be installed (R-273)
The guard owed since Friday morning. Agent v0.128.0 was published as a package and never git-tagged; it was vouched here; and because felhom-host-install.sh fetches an agent's configs from raw/tag/v<version>/configs/, every fresh install and reinstall died at step 5 of 8, as root, on a virgin machine, for most of a day. handleSetArtifacts is the sole UI path to SetArtifactManifest, so the check belongs here and nowhere else. TWO LEGS, because both failed inside two days: the TAG (missing, R-273) and the PACKAGE (pruned from under a still-tagged version, R-287). Either alone catches one of them. It asserts configs/felhom-mkfs-guarded.sh -- the FIRST of the installer's sixteen fetch_raw calls and literally the file whose 404 broke Friday. A test pins the constant, because probing a path that merely exists is how it stayed invisible. The golden gets the package leg only: it has no config tree, so a tag probe would assert something the installer never does. "Could not verify" refuses too, with its own message. No override -- the registry is the operator's own server, so if it is unreachable the vouch can wait. ORDERING IS LOAD-BEARING AND A FAILING TEST FOUND IT. The probes run before resolveArtifactSHA, whose flash conflates "missing", "unreachable" and "bad sha". Probing first means an unreachable registry is reported as unreachable. Five scenarios each naming the wrong outcome; three red-proofs, mutations asserted applied and reverted. With the tag check removed, scenario A reports artifacts_set -- Friday's exact defect returns. |
||
|
|
7264f02172 |
hub v0.101.0 — memoise the artifact dropdown for 60s (R-267, operator ruling)
gates / gates (push) Successful in 28s
v0.100.x removed the serialisation: 26.2s -> ~9.85s mean. What remained was one Gitea package SEARCH per dropdown at 0.20-3.8s depending on load, which concurrency cannot help. Memoised for 60s IN MEMORY. The TTL was ruled by the operator against the workflow that cares: a bake-and-vouch session publishes an artifact and comes straight here to select it, so a minute is short enough not to be noticed and long enough that every reload in that session is instant. NOT persisted. Gitea IS the store for both the version list and the sha; a copy in hub_settings would be a second source of truth that can drift from the registry it describes, and the operator reads the sha here to confirm what they are about to vouch. An in-memory cache dies with the process and can never be mistaken for a record. A failed resolve is NOT cached — a blip must not pin an empty dropdown for a minute. But an empty list from a package that genuinely has no versions IS cached, because 'we found nothing' and 'we could not look' are different answers (CONTEXT S-39, applied to a list instead of a figure). THE FIRST VERSION OF THIS GOT THAT WRONG: the comment said only successful resolves were cached and the code cached the empty list anyway. TestArtifactChoices_FailureIsNotCached caught it before it shipped — which is the argument for writing the test that asserts the comment, and the same class this session spent the day closing. go build/vet/test green, go test -race clean, run separately from this commit. |
||
|
|
e348c4ef6e |
hub: the Gitea client keeps its connections (MaxIdleConnsPerHost was 2)
gates / gates (push) Successful in 28s
Third and last leg, found the same way as the second — by not accepting that the numbers matched the arithmetic when they did not. After the fan-out and the side-by-side resolve the page was ~11.9s mean where ~3s was predicted. Cause: the client used http.DefaultTransport, whose MaxIdleConnsPerHost is 2. Above that Go opens a connection per request and discards it after, so under a 16-way fan-out almost every call paid a fresh TCP setup AND a fresh authentication. Authentication is the expensive half: unauthenticated /api/v1/version answers in ~0.03s while an authenticated package call takes ~0.24s against the same Gitea instance. Transport sized to the fan-out: MaxIdleConnsPerHost 16, MaxConnsPerHost 16 as a ceiling so a large package list can never stampede Gitea harder than the fan-out needs, IdleConnTimeout 90s. |
||
|
|
de0110b8da |
deploy: hub 0.100.1 (the side-by-side dropdown resolve rides this image)
gates / gates (push) Failing after 14m52s
|
||
|
|
9ce8c631d3 |
hub: resolve the two artifact dropdowns side by side, not one after the other
gates / gates (push) Successful in 30s
Follow-up to the fan-out, and the reason for it is worth recording: THE FIRST FIX DID LESS THAN THE ARITHMETIC PREDICTED. Concurrency took the page from 26.2s to ~11-18s, not the ~2s expected, so the gap was chased instead of declared closed. What it found: the slowest single call on the page is not a per-version sha lookup at all, it is the PACKAGE SEARCH (/api/v1/packages/admin?type=generic&q=...), measured in-cluster at 1.1-2.2s each against ~0.24s for a file's metadata. Per-version fan-out cannot touch it — there is one search per package and they ran in series. The two dropdowns are independent, so they now resolve side by side, overlapping both searches and both fan-outs. go test -race clean on the new concurrent paths. Gitea's latency on this box is load-dependent and varies 2-4x between samples, so the CHANGELOG quotes a range rather than a single pair of numbers. |
||
|
|
7855d6355c |
hub v0.100.0 — the Configuration page took 26 seconds, and it was never hashing anything
gates / gates (push) Successful in 21s
MEASURED, NOT GUESSED: GET /configuration -> HTTP 200 in 26.2s. The reasonable guess was that it hashes the artifacts on page load. It does not, and the code already said so: Gitea stores each package file's sha256 and gitea.FileSHA256 reads it as metadata — "a cheap metadata call, the artifact bytes are never downloaded". The cost was never CPU. IT WAS LATENCY x COUNT. artifactChoices made ONE SERIAL round-trip per version, for two packages, capped at 20 each: 2 x (1 version list + 20 sha lookups) = 42 sequential requests at ~0.6s each out through the public ingress. 42 x 0.6 = 26s, which is what the clock said. 1. The sha lookups now run CONCURRENTLY, bounded at 8 in flight. Order preserved by writing into a slot rather than appending — the dropdown is newest-first, and a scrambled sha would show the operator a hash belonging to a DIFFERENT artifact. A failed lookup still drops that version only. 2. The client talks to Gitea IN-CLUSTER (http://gitea.gitea-system.svc.cluster.local:3000, overridable via GITEA_API_URL). Measured from the hub pod: 0.11s against 0.26-1.16s, because the public path adds DNS, the ingress hop and a TLS handshake to each of the 42. Plain HTTP is safe ONLY because it never leaves the cluster network — the registry token rides the Authorization header, so this must not point at a public host without TLS. Unreachable -> the existing graceful degradation to manual text entry, unchanged. DELIBERATELY NOT DONE: caching the sha in the hub's own database. That was the other half of the proposal and it is the wrong shape. Gitea already IS the store; a copy in hub_settings would be a second source of truth that can drift from the registry it describes — and the operator reads exactly this value to confirm what they are about to vouch, so a stale one would be a confident wrong answer. The same reasoning golden_currency_gate.py already records for the vouched version. With the fan-out, a cold load needs no cache to be fast. The cap stays at 20 and now bounds the FAN-OUT too, not just the rendered list. Tests pin order (and that each sha belongs to its own version), per-version failure isolation, and THE CONCURRENCY ITSELF — a wall-clock assertion plus an in-flight counter, so a fast run cannot be luck, and an upper bound so a large package list cannot stampede Gitea. Red-proof: reverting to the serial loop takes 861ms where the concurrent one takes 150ms, and the test fails naming the 26-second page. go build / go vet / go test ./... green (18 packages), run separately from this commit. |
||
|
|
b080ecf411 |
hub v0.99.0 — the hub can see whether the operator can get in (R-260); G-1 gate closes, R-247 closes
oobDegraded tested five things and the sixth never arrived. The agent has emitted `operator_key_configured` on every heartbeat since v0.72.0 — the SAME version that introduced the `oob` stanza carrying it — and store.HostOOBRow mirrored five of the agent's eight OOB fields. With no field for it, encoding/json discarded the fact on arrival, so a box with felhom-sshd active, reachable, a valid config and a configured peer reported `ok` with NO OPERATOR KEY INSTALLED AT ALL. Not a wrong answer: an answer to a question nobody was asking. `operator_peer_configured`, which the hub did read, only says the peer IP is in desired-state — that OOB is MEANT to work, not that entry is possible. Now decoded: operator_key_configured, plus wg_handshake_age_s and healed_at. The last two ride the ALERT TEXT and are deliberately NOT in the predicate — widening a check beyond the fact that is now arriving is how a check stops being read. SCENARIO F, decided on a measurement rather than a preference. operator_key_configured decodes as a POINTER: nil = the agent never said, reported distinctly and never as ok. The version gate was rejected because the field and its stanza shipped in the SAME agent version (v0.72.0), so a stanza without the field cannot come from any released agent; the fleet is 0.113.0/0.127.0 and the vouched floor is 0.127.0. Handled explicitly anyway and pinned, because "cannot happen" is a claim this project has been burned by. THE MESSAGE NAMES THE FAULT. oobDegradedReason is the single source for both predicate and text, so the alert can never name a different fault from the one that fired. The old form derived it separately and had a vocabulary of two — unreachable, or config invalid — with no way to say the key is missing. The operator reads this at 07:00. TESTS DRIVE THE DECODE BOUNDARY. Every hub OOB test before this built a HostOOBRow by hand, and a test written that way CANNOT SEE A FIELD THAT NEVER DECODES — which is how this held a green suite for five weeks. The pre-existing fixture oobReport() also omitted the field, so those scenarios ran against a report shape no released agent produces (same family as R-262). Both fixed. Red-proofs, 8 expected outcomes and 0 wrong, each with the mutation asserted applied: dropping the field returns the false ok; an unconditional check alerts a healthy box; unknown-as-ok restores the silent pass. G-1 CLOSED — scripts/wire_contract_gate.py shipped as ranked, built BEFORE the fixes and seen failing on 40 fields (documentation/tests/wire-contract-gate-2026-08-08/BEFORE.md). Two instrument defects the control caught first: a substring false negative (grep -F healed_at matched privsep_healed_at) and treating dr_recipe as wholly opaque when its top-level sections ARE decoded through an allow-list that already cost offsite_restic (R-122). The prompt for this session said "465 emitted tags, eight unreachable". Checked against the repo: R-260 said "at least eight DECISION-BEARING facts", never eight tags. The real count is 40. R-260 CLOSED (class gated, sharpest instance fixed). R-247 CLOSED (controller v0.209.0). R-264 MINTED and OPEN — the 21 facts with no consumer, allowlisted with reasons so that gating the class could not be mistaken for deciding them. Still open and named: R-246, R-255..R-259, R-261..R-263, and C7's test-comment half. Capability map checked: it claims OOB access is implemented, never monitored, so no row was untrue; what was untrue sat one layer down and the row now records it. repo_gates --fast: all 8 OK. go build/vet/test green in hub, run separately from this commit. |
||
|
|
9657334fb7 |
R-241 FIXED: registers, capability map, STATUS, hub CHANGELOG v0.98.0
gates / gates (push) Successful in 14s
R-241 closed against controller v0.206.0 + hub v0.98.0, following the spike's ruling rather than the obvious reading. The row records what the fix does AND the two real bugs the tests caught rather than review - a missing t.Enabled (caught by an EXISTING test) and a missing falling-edge sync that reintroduced the very defect the epoch exists to fix. R-243 UPDATED, not closed: the STATE it describes can no longer be entered (the mint guard), and what replaces it is VISIBLE rather than silent - the box declares awaiting_recovery_key and the customer is offered the screen. But the ALARM GAP is untouched, for the same three reasons, so a box whose customer never acts still stops backing up with no operator signal. The remaining work is an operator-side signal for a box held past some age, deliberately not bundled into R-241's fix. R-245 NEW - WAITING-ON-OPERATOR, recorded and NOT built: should an undecided customer be auto-abandoned after 30 days? The operator's proposal is recorded WITH the reasoning against it, so the decision can be revisited properly: a reinstall implies a person, so nobody is absent; a customer who cannot find their code gets in touch, which is why the operator LEVERS were the thing worth building; the cost is the customer's own storage allowance; and the real harm is QUOTA, which is a condition, not a calendar. If it is ever built, build it to trigger on the harm with a dated warning, never on a date alone. The capability map's recovery-journey row STAYS FAIL. These are fixes, not a walk - nothing here walked a customer end to end, and the row goes green only when one completes with no operator intervention AND a byte-identical sentinel. R-214, R-202 and R-240 are still open. STATUS compressed rather than extended, per its own one-screen rule, and the "rebuilding throws away the off-site history" line corrected: the cause is fixed, so leaving it as a live defect would be false. hub CHANGELOG v0.98.0 for the superseded-package purge. Highest register ID moves R-244 -> R-245. |
||
|
|
ac4b2a4ba9 |
hub: drop the retained recovery package when the box declares its set-aside history deleted (R-241)
The hub half of the controller's abandonment countdown, and the ONLY reason felhom.eu was touched for R-241 at all. A customer who abandons their old off-site history gets a 14-day countdown. At the end of it the controller deletes the set-aside restic store and then DECLARES offsite.abandon_purge_requested in its report until the retained sealed package that protected that store is gone too. Removing only one half leaves a state that asks a question nobody can answer: a package that opens nothing, or ciphertext nobody can ever decrypt. store.PurgeSupersededEscrowForCustomer is the one place R-198's retention is ever undone, and its doc comment says why that is legitimate here. It NEVER touches host_escrow - the current package covers the key the box is using now and is what makes its live backups recoverable. Only host_escrow_superseded rows go. The handler acts on the box's DECLARATION, never an inference, on the same principle as offsite.state: the hub cannot see that a remote store was deleted and the box can. It is placed immediately BEFORE the ACK is built, deliberately. GetEscrowStatusForCustomer is read after it runs, so the SAME response that carries the request's effect is what closes the box's two-phase commit - no second round-trip, and no window in which the box believes it is still owed. The declaration repeats on every report until that ACK stops reporting a superseded package, so a lost request retries by itself. A purge failure is logged at ERROR and never swallowed: the box keeps declaring, so it retries, but an operator must be able to see that the two halves are apart right now. An idempotent re-declaration (already purged, the box has not yet seen the confirming ACK) logs at DEBUG and is not an error. Audit event offsite_abandon_purged is hub-internal, like the pbsdr_* and offsite_selfheal_* events - allowedEventTypes governs the box-pushed POST /event surface, not this. No agent change. No deletion has been performed against any real store. Green: go build, go vet, go test ./... all clean in hub/; repo gates OK. |
||
|
|
04ac465da6 |
CAMPAIGN-11 Phases 2+4: the fault journal, the campaign document, and hub v0.97.1's missing heading
gates / gates (push) Successful in 8s
Phase 2 (eleven injected faults) and the §4 positives that were owed. - §4.1 MEASURED, twice: the box's rendered GetFloor() is 0.200.0, and a cold-started controller logs "settle-gate: GO — at/above floor 0.200.0" against the same line reading "floor still unknown" while the hold was in force. Also corrects the brief's plan: SetFloor's line is u.dbg(), gated on logging.level=debug and written to the logger, so it can NEVER reach the debug ring — a restart alone would not have produced it. - §4.2 still NOT measured, deliberately: the venue has an off-site target, so needsOffsiteCredential correctly returns false. Recorded, not inferred from the unit test. - F1 PARTIAL, F2 PASS, F3 FAIL, F4 FAIL, F5 PASS, F6 PASS, F8 PARTIAL. F3+F4: a hub outage and a stopped agent are both rendered as "this code does not open your package", in 0.056 s and 0.030 s — no unseal attempted. The agent's own err field distinguishes them exactly and it is discarded at the HTTP boundary; the R-216 capability gate answers source=version and cannot see reachability. - R-217's fix HOLDS under exactly its fault (F5), verified with the false-claim strings absent and accented positive controls present. hub/CHANGELOG.md: v0.97.1 had no heading of its own — the change was written into the v0.97.0 entry while the deployed tag is 0.97.1. Given its own entry, marked as added retroactively. Second occurrence of the class (agent 0.90.1). Evidence: documentation/tests/campaign11-evidence-2026-08-05/journal-phase24.md No product code changed. |
||
|
|
a7f1d277b1 |
hub: the held-floor REASON must match the CAUSE (CAMPAIGN-11 follow-on)
gates / gates (push) Successful in 7s
v0.97.0 introduced a second hold reason and left both surfaces printing the first. The freshly deployed hub logged, for the campaign box: managed floor HELD for c11: agent "0.125.0" < MinAgent 0.113.0 which is FALSE — 0.125.0 is ABOVE 0.113.0. That box is held because its floor sits above the vouched golden, not because of its agent. CLAUDE.md's corollary exactly: when a verdict changes which field it counts from, the alarm text has to change with it, or a true alarm reads as one to dismiss. Both the ACK log line and the Hosts-dashboard HeldReason now come from one ManagedFloorDecision.HoldReason(), and TestResolveManagedFloor_HoldReasonMatchesTheCause pins each reason to its cause. |
||
|
|
7e1d2898bd |
hub v0.97.0 — the floor stops being served past the agent it depends on (CAMPAIGN-11)
gates / gates (push) Successful in 8s
R-216, the hub half. ResolveManagedFloor's own comment says it exists to "never push a controller past the agent it depends on", and it compared against ArtifactManifest.MinAgent — which by ITS own comment describes the GOLDEN's controller. publish-train-rules.md rule 3 states the rule about the FLOOR's controller. Measured live: golden 0.192.0 / MinAgent 0.113.0, floor 0.200.0, agent 0.120.0 — served, and the box was pushed onto a controller needing agent 0.125.0. A floor ABOVE the vouched golden is now HELD with its own reason (HeldBeyondGolden), reusing Part D's dashboard visibility. Nobody types a number twice: the vouched MinAgent keeps its meaning, the guard stops applying it to versions it does not describe. An uncoupled release is untouched; an unparseable golden degrades rather than gating. R-222: the report ACK's escrow object gains superseded_present / superseded_at, counting only rows that actually carry an identity blob. One boolean and one timestamp, for one message. No read path — that link is still unbuilt. Red-proof: removing the floor-above-golden branch reproduces the campaign's measurement. |
||
|
|
b2462b8f36 |
CHANGELOG: hub v0.96.0 (R-204 item 4, hub half)
gates / gates (push) Successful in 7s
|
||
|
|
f62a115891 |
R-204 item 4 (hub half): the hub answers a rebuilt box's request (hub v0.96.0)
New internal/offsiteheal, the sibling of pbsdrheal: it acts ONLY on the state the box declares, sustained across two distinct reports, re-staging the stored credential before ever minting a new one. A healthy box is a pure no-op; it never blind-timer-reissues and never re-runs a provisioning step. RESTAGE IS POSSIBLE because the stored value survives a consume — established from the schema and ConsumeOneTimeSecret (which stamps consumed_at and nothing else), not inherited from the PBS analogy, and pinned by a test that asserts the SAME value comes back. reportHasOffsite is TIGHTENED to require enabled:true. Its comment asserted that presence == applied-on-the-box, and the declaration deliberately breaks that premise; left alone it would have read a request for help as proof the tier was applied. Provably a no-op for every report shape that existed before, because an attached object has always carried enabled:true. R-192's guard half is CLOSED BY REPLACEMENT: the delivery checker's counting inference read the OLDEST 500 reports after a consume — all predating a rebuild, which is why demo-hp sat stranded for 108 reports under a confident regressed-shape verdict. A declaration outranks both inferred shapes, and the checker stands down with a record so the two mechanisms cannot double-issue. No escrow ceremony is ever run or requested: credential automatic, key customer-present. |
||
|
|
5c7d67102d |
CHANGELOG: hub v0.95.0 (R-196 / R-204 item 2)
gates / gates (push) Successful in 7s
|
||
|
|
d1a8edb332 |
R-196 / R-204 item 2: a re-issue no longer marks a healthy escrow stale (hub v0.95.0)
ReissueCredentials marked the escrow stale on every re-issue, on precautionary grounds — the box's re-apply MIGHT mint a fresh repository password. It usually does not. A stale flag withholds restic_pw_sha256 from the ACK, which stops the controller's auto-confirm, which leaves EscrowState pending, which makes OffboxRunnable false: every off-site backup refused on a box whose key was never in doubt — and the customer told to re-run the one ceremony that would have superseded the key just recovered. The case it guessed at is measured elsewhere: the controller's Scenario-F re-check compares the sealed hash against the live repo password on every ACK (and the mark was BLINDING it by emptying that hash), and R-197's offsite_repo_key_changed fires on a proven difference across a supersession. offsite_reissued is unchanged. MarkEscrowStale is kept without a caller so a future EVIDENTIAL writer has the mechanism, with a test pinning it live. TestReissue_InvalidatesEscrow is replaced by its exact inverse. |
||
|
|
435f4a5229 |
hub v0.94.0: a box can fetch its own sealed recovery package (R-199 link 6)
gates / gates (push) Successful in 7s
Link 6 of the recovery chain had no client. The hub has served the identity blob since
slice 10D from handleReEnroll / handleGetRestoreDirective, gated on operator-armed recovery
mode and the global key -- and nothing in the agent, the hub UI, any script or any runbook
ever called either. The only documented retrieval was sqlite3 writefile() by hand on a
kubectl cp-ed database.
GET /api/v1/hosts/{host_id}/escrow is the box-authenticated mirror of the PUT that put the
blob there. Self-scoped (a per-host key reads only its own; global may read any). A host with
no bundle gets 200 {present:false} -- a 404 is indistinguishable from an unknown host and a
bare empty 200 from a zero-length blob.
THE TRADE IS RECORDED IN THE HANDLER, not inferred: obtaining the blob used to require the
operator to arm recovery mode; now whoever controls a rebuilt box can obtain it with that
box's own credential. They still cannot open it -- the hub has never held R and a wrong code
fails closed at age's scrypt KDF. The mitigation is that every retrieval raises
escrow_blob_served (warning, operator-only), recorded before the bytes leave.
escrowSelfServiceRetrieval is the single decision point: flip it to false and the endpoint
additionally requires recovery mode, changing nothing else.
The operator-driven DR path is untouched, pinned by a test. Red-proofs observed: removing the
ownership check serves host B's blob to host A; removing the record makes it silent.
|
||
|
|
91cabdde1b |
hub v0.93.0: the retention keeps the key it was built to keep (R-198) + three honesty fixes (R-197, R-192, R-196)
gates / gates (push) Successful in 7s
R-198 — host_escrow_superseded shipped with `blob` (the K-escrow / PBS datastore key) and
identity_blob was added to host_escrow LATER, never here. The offsite restic REPOSITORY
password lives in identity_blob. So demoteCurrentEscrowTx -- whose own comment calls it "THE
ONE escrow row-copy routine" -- retained the whole-guest key and silently dropped the off-site
data key, which is the secret the retention was built to preserve. And because the copy happens
as the new blob overwrites the old, the destroying act was the ESCROW CEREMONY: the exact thing
a rebuilt box tells its customer to run, on a card promising in Hungarian that the old backups
stay recoverable. Both demo boxes crossed that line on 2026-08-04.
- identity_blob added to the table (CREATE + additive ALTER) and carried in the shared copy
routine, so BOTH callers are fixed at once: re-escrow and host-delete demotion.
- ListSupersededEscrow reads it back; store.HostEscrow gains IdentityBlob.
- CountCurrentEscrowWithIdentity is the census of who the fix protects.
- Nothing is backfillable: pre-v0.93.0 retained rows have no blob and their sources are gone.
- Tests assert the CONSEQUENCE (a retained row can still yield a repo password), which is why
the pre-existing retention test stayed green for two months asserting the mechanism.
R-197 — SaveHostEscrow returns the hash it replaced; the escrow PUT raises
offsite_repo_key_changed (warning, operator-only, edge-triggered) when both hashes are known and
differ. No hash value travels. Severity chosen for the world v0.93.0 creates: with the identity
blob retained, a changed key is "this history now depends on an older recovery code", not a loss.
R-192 (half) — the stuck alert now reports the two shapes it actually covers, burned and
regressed, each stating its own measurement; the regressed text withdraws the Re-issue
recommendation. Every self-heal refusal leaves a notification_log row with its reason. The
guard's logic is unchanged; its 500-oldest-reports scoping stays OPEN and the window is named in
the alert text so the limitation travels with the number. offsite_delivery_stuck and
offsite_credential_restaged are added to operatorOnlyEvents -- neither was registered and neither
has a customerMessages entry, which is not a block.
R-196 — five comments (not the three the spec expected) claimed ReissueCredentials rotates the
restic repo password. It resets the PROVIDER password and cannot touch the repo password, which
is generated on the box. All five corrected; the staleness mark documented as precautionary. The
BEHAVIOUR stays open.
Not in this release: R-199, R-200, R-201 remain open -- the chain that hands the key back is
still unassembled. Part 5 hit its gate; the orphan card is untouched (R-202).
|
||
|
|
7fff45d688 |
R-195: a customer with no machine ever bound does not alarm (hub v0.92.0) + R-193/R-192 spike
gates / gates (push) Successful in 7s
Part 4 (ships): `david` — a prospective customer with hosts=0, host_deletions=0, reports=0 — e-mailed an expected_dbdump_missed ERROR at 03:00 UTC three mornings running. The existing down-skip could never cover it: it reads the staleness checker's state, which is seeded from a query over the `reports` table, so a customer that never reported has no state at all and GetState() returns "" rather than "down". store.HasEverBoundHost (hosts row OR host_deletions tombstone) is consulted once per customer at the top of the deadline loop. The discriminator is "was a host EVER bound", never "has a report arrived" — a box installed and never heard from is a real fault and keeps alarming. Fail-OPEN on a read error. Red-proof observed: removing the guard fails with `got [expected_dbdump_missed]`, verbatim the event david sent. Parts 0-3 (spike, NO production code for R-193/R-192): audits/SPIKE-offsite-credential-recovery-2026-08-04.md establishes that the one-shot provider password is the RECOVERABLE secret and the restic repository password is the irreplaceable one — and that a guest rebuild mints a fresh one, orphaning the previous off-site history. Measured without touching a box, by comparing host_escrow.restic_pw_sha256 against host_escrow_superseded: BOTH demo boxes changed (demo-hp 15 snapshots / 40.9 MB, demo-felhom 36 snapshots / 1.14 GB). demo-felhom's "lucky" 76-second recovery restored delivery and not the repository, silently, for 13h. ReissueCredentials does NOT rotate the restic password (R-39's record and two hub comments are wrong -> R-196); candidate (b) is not implementable against a zero-knowledge escrow; candidate (a) already exists as F3 and is wired to the wrong event. Ends in ranked options and an unanswered question for the operator. R-195 SHIPPED; R-196 + R-197 filed; R-192 + R-193 updated, neither closed. |
||
|
|
046df303b6 |
hub v0.91.1 — observation may only WIDEN a tier's window, never tighten it (R-86)
gates / gates (push) Successful in 7s
Found by checking v0.91.0 against the live box, not by review. demo-felhom's two retained PBS snapshots sit 8h54m apart (one is a healing artefact), so the mean-gap estimator reads a WEEKLY tier as nine-hourly: x4 = 36h, the 7-day floor lifts it to 168h, and a weekly tier proved weekly reaches ~8.25d of proof age. The false alarm this task exists to prevent would have returned within a week, on the box it had just shipped to. restoreProvenWindow now takes max(observed, declared). A gap SHORTER than the declared rhythm is routine and means nothing (a retry, a manual run, a heal, a catch-up); a gap LONGER than it is real information. Cost stated: a tier running faster than its declared rhythm gets a slower stale signal — the right direction for a signal that means 'unverified', since 'broken now' is a different event. |
||
|
|
323f45a5ef |
hub v0.91.0 — the staleness window learns each tier's own rhythm (R-86 Part 2)
gates / gates (push) Successful in 7s
Ships WITH agent v0.121.0, not after it. The agent now proves a tier once per
ARCHIVE GENERATION, so a weekly tier is proved weekly — in perfect health. The
flat 7-day restoreProvenStaleAfter derived its number from the 24h cadence R-86
removes, and a healthy weekly tier's proof age reaches EXACTLY 168h just before
its next proof: it sat ON the line, so any ordinary delay tipped it into a
nightly alarm about a working system.
restoreProvenWindow(tier, observed, ok):
- the tier's own archive interval, OBSERVED from reports the hub already holds
(pbs_snapshots + successful backups attributed by TARGET TYPE, slice A.4)
- x4 generations = the same tolerance the flat constant expressed
- floored at 7d (never tighter than before), capped at 12d (strictly inside the
2-week offsite retention)
- falls back to the DECLARED rhythm (26h host / 8d offsite — the thresholds the
backup-freshness checker already uses) when history is too short to observe
one; falling back to the FLOOR would recreate the false alarm on a fresh box
Kept: absence is UNKNOWN until the anchored window passes; the signal stays
edge-triggered; failed and stale remain distinct events. Every reason string now
states the window it was judged against (R-100's corollary).
Also backfills the missing v0.90.1 CHANGELOG entry (deployed since
|
||
|
|
f21e7caed1 |
hub v0.90.1 — the digest's per-app lines stop repeating the filesystem figures (R-182)
gates / gates (push) Successful in 7s
Found by reading the first REAL digest, not by design. Every app row ended with the same usage clause the mail already prints once on its own Filesystem line. On a two-app box that is untidy; down a list of a dozen it is the same forty characters twelve times, pushing the part that DIFFERS off a phone screen at 07:00 — the only moment this mail has to work. The reserve's refusal message is authored for a single-app alert where naming the filesystem is right, so the message is unchanged; the digest trims the duplicate when rendering. trimRepeatedUsage removes ONLY an exact "— <target path>:" suffix, so an unrelated reason is untouched and a reason that is nothing but the usage clause is left alone rather than emptied. Also updates TestRecoveryUnitCaptureFailed_NeverReachesTheCustomer, which required the OPERATOR to be emailed a per-app capture failure. That was correct when the event was the only signal and is wrong now that it is the record and the digest is the notification. Its customer-safety claim is unchanged and is why the test still exists; the operator assertion is inverted with the reasoning written in place, and R-158's guarantee is shown to have MOVED, not weakened. |
||
|
|
dd40f85bb8 |
hub v0.90.0 — a dropped notification leaves a trace, and the backup digest arrives (R-182)
gates / gates (push) Successful in 7s
processOperator's cooldown no longer returns bare. It dropped the event BEFORE LogNotification, so a suppressed operator alert and an event that never happened were indistinguishable — from the operator's side and from the hub's own records. Measured 2026-08-03: nine recovery_unit_capture_failed events arrived, two were mailed, seven left no row anywhere. That is why the defect took a day to get the right way round: there was nothing to read. A suppressed operator event now writes a `suppressed` row carrying the message and the key that suppressed it. This applies to EVERY operator event, not only the one that exposed it. It does NOT change the cooldown's duration or semantics. backup_run_failures: the per-run digest. In allowedEventTypes AND in operatorOnlyEvents — allowlisting alone does not make an event operator-only, and FormatCustomerEmail falls back to the raw English message rather than blocking. A test demonstrates a customer with the type enabled receiving nothing. recordOnlyEvents: a third routing class — stored and recorded, never mailed. recovery_unit_capture_failed moves here: it is the record, the digest is the notification. A register rather than downgrading severity to info, which would relabel a genuine failure as informational everywhere it is queried. cooldownRunSuffix: a sibling of cooldownTierSuffix, not a branch inside it, so tier keeps byte-identical semantics and R-97a's tests are untouched. It makes the cooldown effectively inert for the digest, which is the intent — a digest is already rate-limited by construction; the refresh sweep sends no run_id and so stays under the ordinary hourly cooldown. The email renders as a list, not a JSON blob. An absent space reading renders as unavailable, never as zeros. |