fcffaf573a1ba5be943796c28cfb3127ac21f4c3
911 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
fcffaf573a |
v0.209.0 — R-247: the box stops saying a false thing about its own recovery package
gates / gates (push) Successful in 17s
The answer was on the wire and was discarded at the boundary, for the third time. The hub has sent `escrow_stale` in the report ACK since v0.57.0 (json:"escrow_stale,omitempty"). report.EscrowStatus had no field for it, so encoding/json dropped it, and an empty restic_pw_sha256 had exactly one possible reading here: "hash-less supersession". On demo-hp that reading was false in EVERY clause for four days, and the box told the customer so in its own words. The hub HAD the hash and was withholding it because the escrow row carries a stale flag (R-246); there had been no supersession; and the bundle DID cover the password — the hashes matched exactly. Fixed by receiving the field. EscrowStatus.Stale decodes, and reconcileEscrowed tells the two conditions apart: a withheld hash now reports that the hub has flagged the row and is withholding, that this box therefore cannot verify its bundle either way, and that it is NOT established that the bundle fails to cover the password. The genuinely hash-less case keeps its wording. Deliberately NOT changed, and said rather than skipped: the stale verdict itself (the hub's flag is still the hub's verdict; runs still continue), and the customer-facing Hungarian card copy. Clearing the wrong flag is an operator act hub-side (R-246); re-wording the card is UI work with its own review path. This change is the wire and the diagnosis. Found by felhom.eu/scripts/wire_contract_gate.py (G-1), which was built first and seen failing on 40 fields before anything was fixed, and which now refuses any new field of this shape. go build / go vet / go test ./... green, run separately from this commit. |
||
|
|
37b5ba08a7 |
REPORT: v0.208.0 — both R-254 sites, the guard's measured holes, and what the live read could not prove
gates / gates (push) Successful in 13s
Records the deliverables, and is explicit about the limit on the live half: the curl of an app info page could not be done, and names exactly what was tried — crafty-controller is the only app declaring initial_credentials and is deployed nowhere, and demo-hp's dashboard password in ~/.config/credentials no longer authenticates (200 with no session cookie). A probe of the new routes was discarded because its control killed it: real and bogus paths both 302 behind the auth middleware. §7.2's answer including the part that contradicts the task's premise: no line in the repo says 'no silent auto-fill'; the rule is CONTEXT.md:2070 about accidental EMPTY-password deployments. The hidden input is deliberate and untouched. §7.4's measurement: the gate covers all 36 templates and catches a launder through a local variable, but is blind to a secret under a neutral page-data key — the exact shape of site two. Runtime coverage is 4 of 27 pages. Filed as R-255 rather than described as complete. §7.3: no evidence of actual exposure on the fleet, with the limit stated — it is a current-state measurement and nothing recorded reads, which was part of the fault. Also corrects v0.207.0's report: html/template STRIPS HTML comments; they do not ship in the response body. Measured. |
||
|
|
27d1165962 |
v0.208.0 — R-254: the last two secrets leave the page source, plus a gate against a fourth
gates / gates (push) Successful in 17s
Site one. app_info.html rendered {{.InitialCreds.Password}} into a hidden span —
a REAL per-install credential, read live out of the running container, in the
response body of every render. The page now carries the non-secret half plus a
boolean; the value comes from POST /apps/<slug>/initial-credentials/reveal, which
RE-READS the container rather than serving a cached copy (caching it in the
handler would put it back in the body one layer in). no-store, CSRF-covered,
logged as an act. Both buttons go through it. A reveal that cannot read the value
SAYS SO rather than returning an empty string that renders as a blank password.
Site two, established before changing. The hidden input is NOT the defect and was
left alone: it fires only pre-deploy, and README §318 documents why the value must
round-trip — the customer notes the generated secrets down and submitting them
back is what makes the saved value the same one they saw. The defect was the
neighbouring READONLY input, which on an ALREADY-DEPLOYED app rendered the secret
into a page with nothing to submit. Fixed by POST /stacks/<name>/auto-field/reveal,
authorised by requiring a type:secret auto-field of that stack. Both directions
pinned.
The premise that this contradicted a repo rule does not hold: the rule is
CONTEXT.md:2070 'Password fields require explicit input — prevents accidental
empty-password deployments', about EMPTINESS. No line in the repo says 'no silent
auto-fill'.
The gate. scripts/secret_in_markup_gate.py, registered in controller_gates.py,
convicts any template expression that names a secret unless allowlisted with a
reason. Its limits are MEASURED and in its docstring: it catches a launder through
a local variable (the assignment names the secret) but is blind to a secret
arriving under a neutral page-data key — verified both ways. That is the shape of
site two, which this gate would NOT have caught. The runtime body assertion covers
all shapes but only 4 of 27 page templates; the other 23 are R-255, filed rather
than glossed. Two nets, different holes, both named.
Correction to v0.207.0's report: HTML comments do NOT ship in the response body
here — html/template strips them, text/template does not. Measured. A red-proof
planting a secret in a comment therefore correctly does not fail.
|
||
|
|
62998aab4f |
REPORT: v0.207.0 — R-249 before/after on a live box, the census, and what was NOT proven live
gates / gates (push) Successful in 18s
Records the deliverables: the raw response body before (1 occurrence, v0.206.0) and after (0, v0.207.0) with a positive control in both directions; the §7.1 census finding two more instances of the render-then-hide pattern (R-254, one a real per-install secret); §7.2's decision and why the promise was the wrong half; every changed Hungarian string; all eight tests with their red-proof outcomes. States plainly what was NOT proven live: R-252/R-253's notices could not be rendered on VM 325 because both states are rebuild-only and the box re-registers a drive on restart — the live run therefore exercised Scenario E instead, and the notices are pinned at the template + predicate level with red-proofs. Also records that red-proof D caught a fault in my own work: the explanatory HTML comment quoted the old sentence, and HTML comments ship in the response body, so the contradiction was still on the page and the assertion forbidding it could never fail. |
||
|
|
8dbbc98ff2 |
v0.207.0 — R-249: the retrieval passphrase leaves the page body; R-252/R-253: two refusals learn to say what to do
gates / gates (push) Successful in 18s
R-249. settings_security.html rendered the passphrase into a display:none span behind a Megjelenit button. That toggle stops a browser DRAWING the value and nothing else — the plaintext was in the response body of every render, so a curl of the page returned it. Found by exactly that: it landed in a session transcript while driving the documented rebuild path. The codebase already stated this rule for the recovery code and this page did not follow it (escrow_handlers.go: 'reveal (claim XHR only — R is NEVER templated server-side into HTML)'). The page now carries only HasRetrievalPassword; the value comes from POST /settings/retrieval-password/reveal — CSRF-covered because POST, no-store, and LOGGED as an act, which reading it off the markup never was. The tests assert the RAW RESPONSE BODY. Every test that asked what the customer sees passed while the bytes carried the secret; that is why this survived. Census: the render-then-hide pattern appears twice more — app_info.html (a real per-install app password in a hidden span) and deploy.html. Filed as R-254, NOT fixed here. R-252. A rebuilt box keeps its drives but loses their REGISTRATION. The restore page now states that before the customer presses anything, says the backups and drives are both still there, and links to Tarhely > Meghajtok. Page and resolver ask ONE question — HasRestoreDestination() reads the same GetSchedulableStoragePaths() the scratch resolver reads. R-253. The list promised 'a visszaallitas elobb ujratelepiti' three lines above a refusal that fired BECAUSE the app was not installed. The promise was the wrong half: reconstitution writes to the app's own GetStackHDDPath, which exists only once the CUSTOMER has chosen a drive at deploy time. Auto-reinstalling would mean the product making that choice for them. Copy now says to install first and routes to /stacks/<app>/deploy. Both notices are conditional — a healthy box renders as before, pinned by a test that fails if either becomes unconditional. |
||
|
|
3d3b4496f3 |
REPORT.md — R-241 fixed, deployed, live-validated on both demo boxes
gates / gates (push) Successful in 22s
Scenario A's live result first: on demo-hp in the rebuilt shape, no key was minted on the real start-up offsite-apply path, and the hub received the state it reports instead - offsite.state=awaiting_recovery_key with enabled:false. Key restored byte-identical afterwards. Includes Q4's seven rows mapped to the three states, the SEC 7.2 choice and why, SEC 7.3's answer on the new-code button, every changed Hungarian string quoted, all nine red-proofs with what was mutated, the R-245 reasoning, and three observations noticed but not acted on. |
||
|
|
0a9158d53e |
docs for v0.206.0: CHANGELOG, CONTEXT, REUSE, README
gates / gates (push) Successful in 19s
CHANGELOG v0.206.0 with the ruling that reversed the fix, the three changes,
the SEC 7.2 staleness decision, Q7's closed trap, and the two bugs the tests
caught rather than review.
CONTEXT carries the three rules this session established, in the form the next
session needs them:
- a box does not create a repository key while the hub holds a sealed
package for it;
- the fact that answers a question must be kept where the question is asked;
- fix the state, do not remember that it is wrong.
REUSE gains four rows, each carrying the trap rather than just the signature:
the mint guard is a CONJUNCTION and t.Enabled is load-bearing in the derived
predicate; the discriminator ships INERT unless wired in main.go's confirmer
literal; the countdown removes BOTH halves or neither and must be driven by an
injected clock; and the epoch must be synced FIRST and unconditionally or the
falling edge is lost.
README documents the three customer-visible changes and the operator levers.
No version literal was edited: the controller version is ldflags-only.
|
||
|
|
72368654e4 |
R-241 part 5: escalating reminders, and operator levers for a running countdown
REMINDERS (SEC 2.3). The offer epoch now stamps when it began, and the undecided reminder escalates in EMPHASIS at 1, 3, 7 and 14 days. THE READING IS STATED BECAUSE THE SPEC IS AMBIGUOUS, and it is written into the code where it can be corrected. For an ABANDONING box, 5/3/1 are unambiguously days REMAINING before a deletion. An undecided box has no deadline - nothing counts down to anything, because SEC 7.5 deliberately does NOT auto-abandon - so 14/7/3/1 cannot be "remaining" and are taken as days ELAPSED, with the wording firming up rather than the bar appearing and disappearing. If the operator meant something else, one function changes. The stamp is re-set on every entry into the offered state, so a box that settles and is later rebuilt starts its ladder again instead of inheriting an old one. OPERATOR LEVERS (SEC 7.5). --abandon-status, --abandon-extend=N and --abandon-stop on the controller CLI, beside the existing operator subcommands. They exist because the path that ACTUALLY happens is the customer telephoning, and support needs something to press. They live on the CLI and not in the customer UI deliberately: extending a deletion the customer asked for is an operator judgement, and a customer who wants it stopped already has the self-service route - they recover with their code, which cancels it. BOTH REFUSE RATHER THAN NO-OP, in two situations: when no countdown is running, and when the store has already been deleted. A silent success is the thing an operator most easily mistakes for "handled" - they would tell the customer their data was safe when it is gone. Pinned by two tests. --abandon-extend counts from NOW, not from the old due date, and a test proves the old date passes without deleting anything. Green: go build, go vet, go test ./... all pass; controller gates OK. |
||
|
|
de39e47f53 |
R-241 part 4: the three-state surface, and the copy tells the truth about the date
FULL PAGE ONCE PER ENTRY, NOT ONCE EVER. "Most nem" used to set a flag that
nothing ever cleared, so a box that abandoned its history and was rebuilt
months later - a genuinely NEW situation - would never see the page again. The
offer now carries an EPOCH, advanced on the edge into the offered state, and a
dismissal is recorded against the epoch it was made in. A fresh entry passes
the dismissal by arithmetic, with nothing to clear and nothing that can be
forgotten to clear.
That is NOT the flag the operator's ruling forbids. The forbidden thing
remembers that the customer decided so the screen can be suppressed while the
state stays wrong. This records WHICH SITUATION a dismissal was about.
A REAL BUG, caught by the test and not by review: the first draft returned
early from recoveryInterrupts when the offer was false, so the FALLING edge
was never recorded, RecoveryOfferActive stayed true through a settled period,
and the next entry counted as a continuation. The page never came back - the
exact defect the epoch exists to fix, reintroduced inside the fix. The sync is
now unconditional and the ordering is commented as load-bearing.
THREE LEVERS, THREE SCOPES, and none of them removes the route:
- clicking the bar away -> a browser SESSION cookie, cleared on login, so
the reminder is genuinely back at the next login. Nothing persisted.
- "ne emlekeztessen ujra" -> durable, epoch-scoped, silences the BANNER ONLY.
It starts no countdown, abandons nothing, and a fresh entry reminds again.
- "most nem" -> suppresses the full page only, as before.
The entry point on /backups/remote is bound to the OFFER and to nothing else,
pinned by a test that fires all three dismissals and asserts it survives.
SEC 7.3 / Q7 - THE TRAP DOES NOT SURVIVE THIS SESSION. While a recovery is
outstanding the "Helyrealitasi kod letrehozasa" button is UNAVAILABLE, not
merely captioned: creating a new code seals the current key, demotes the
package that opens the earlier history to retained custody that no shipped
path can read (R-199), and re-enables the recovery screen through the orphan
route while invalidating the code that screen accepts. A warning beside a
button is a warning people click past. The card now explains and points at
/recovery instead.
SEC 2.4 - the abandon confirmation changes with the behaviour. It used to
promise "felretesszuk - nem toroljuk". It now states the grace in days (from
the constant the countdown actually uses, never a literal in prose), that the
sealed package goes with it, that the customer can change their mind, where
the date is visible, and that the question does not come back afterwards.
The countdown is shown on /backups/remote for the WHOLE window - the bar
elsewhere is a nudge, this is the record, and a deletion date must be findable
on a quiet day too.
Tests: once-per-entry across a full settle-and-re-enter cycle; the banner
dismissal proven to be a session cookie (MaxAge 0, no Expires) and to persist
nothing; the opt-out proven to silence the banner while leaving the offer, the
route and the countdown untouched, and to remind again on a fresh entry; the
entry point surviving all three dismissals; a settled box showing nothing; and
the back-redirect refusing "//evil.example".
An existing test (TestRecovery_E) was updated: it asserted the legacy boolean,
which the epoch replaces. It now asserts the dismissal landed on the current
epoch, which is the stronger property.
Green: go build, go vet, go test ./... all pass; controller gates OK.
|
||
|
|
a5d90ff801 |
R-241 part 3: abandoning starts a 14-day countdown that ends the question
Until now "set aside" renamed the remote store and touched neither the escrow
nor the key, so the hub went on holding a sealed package for a key the box no
longer used. Shape (c) compares those two, finds them different, and offers
recovery - correctly, and for ever. A customer who had already said "I do not
want the old data" would be asked again at every login.
The operator's ruling is that the answer is NOT a "they decided" flag: fix the
state, do not remember that it is wrong. So the decision starts a countdown,
at the end of which the set-aside store and the sealed package that protects
it are removed TOGETHER. Afterwards shape (c) has nothing to compare and the
offer falls silent on its own - because the state is right, not because
something remembers it once was not.
THE GRACE IS REAL. The recovery offer stays reachable for the whole 14 days;
that is the change-of-mind path, and a grace in which recovery is impossible
would be decorative.
BOTH HALVES OR NEITHER. Removing only the store leaves a package that opens
nothing; removing only the package leaves ciphertext nobody can ever decrypt.
The two cannot be atomic across two machines, so it is a two-phase commit:
delete the store, record a durable marker, and keep DECLARING
offsite.abandon_purge_requested until the hub's ACK stops reporting a
superseded package. A crash between the halves re-declares on the next sweep;
it never leaves the pair half-removed and silent.
HUB HALF - SEC 8.2 ANSWERED: yes, the hub was needed, and only for this.
store.PurgeSupersededEscrowForCustomer is the one place R-198's retention is
ever undone, and it never touches host_escrow (the package covering the key
the box uses now). The handler acts on the DECLARATION, never an inference,
and is placed immediately BEFORE the ACK is built - so
GetEscrowStatusForCustomer reads the effect and the SAME response closes the
box's two-phase commit. No second round-trip and no window where the box
thinks it is still owed. felhom-agent was NOT touched.
The countdown starts in ResetOrphanedRepo, NOT in the shared helper: the
helper is also the unclaimed auto-reset path, where nobody decided anything,
and an as-delivered box tidying a stranger's leftover store must not get a
customer's deletion clock. Pinned by a test.
Cancellation is wired into the recovery unlock, BEFORE the tier-up and the
listing - those can fail, and a countdown surviving a successful unlock
because a later step errored would delete the history the customer just
proved they can open.
The sweep is a Daily job at 05:10, not on the backup leg: it must run on a box
whose tier is not configured for runs. Quiet by construction on every box with
no countdown, and that silence is asserted.
Tests (all clock-injected; SEC 7.4 forbids shortening a live timer):
Scenario E (aside + package kept + countdown + offer still reachable, and
NOTHING deleted), Scenario F (both halves, the declaration repeating, the
close-out), Scenario G (cancel, path still nameable, no later deletion),
plus: not closed out while the package remains, a transport failure leaves the
countdown due and retrying, the no-op sweep issues zero remote commands, and
the unclaimed auto-reset starts no countdown.
RED-PROOFS, each with the mutation confirmed present in the file first:
F1) store deletion skipped -> Scenario F FAILS (no rm issued)
F2) declaration dropped from the report -> Scenario F FAILS (the hub is
never asked; the package would outlive the store for ever)
G) CancelAbandon made a no-op -> Scenario G FAILS (uncancellable countdown)
Green: controller and hub both build, vet and test clean; controller gates OK.
NOTHING WAS DELETED ANYWHERE - the terminal step has only ever run against
in-test fakes.
|
||
|
|
a491abef6c |
R-241 part 2: the comparison the box already makes becomes the thing that offers recovery
THE FACT WAS COMPUTED EVERY CYCLE AND KEPT NOWHERE. EscrowAutoConfirmer.Reconcile
has compared the hub's restic_pw_sha256 against the local key on every ACK since
SLICE 3. On the final-walk venue it logged, at 03:28:03Z and thirty-five minutes
before the customer looked, "the hub's escrow blob does not cover the CURRENT repo
password (hub hash 30ef574f != local 9b4a9a9d)" - and dropped it. The recovery
screen, evaluating in the same process, went on asking a question that could not
see it.
Now persisted: settings.HubEscrowKeySHA256 + HubEscrowKeyCheckedAt, recorded
UNCONDITIONALLY in Reconcile beside RecordPresence and RecordSuperseded - same
place, same reason: the box that needs it most is the rebuilt one with no target,
on which every gate below returns early.
OffsiteRecoveryOffer gains SHAPE (c): the hub holds a package for a key OTHER than
the one we are using. (a) and (b) are both proxies for that question and both have
now been wrong in opposite directions - (a) goes false the moment anything mints,
(b) is unreachable while the escrow is pending.
SEC 7.2, decided deliberately and stated in the code:
- a KNOWN DIFFERENCE offers, however old the reading. Age is not gated on. Both
sides are local; only the hub's half can be stale, and what the hub holds does
not change without a ceremony THIS box runs, which refreshes the hash on the
next ACK. Gating on age would make a box offline from the hub silently stop
offering - the exact failure this session removes. CheckedAt is persisted for
diagnosis, not as a gate.
- an ABSENT hash falls back to (a)/(b) and does NOT offer. "" is the hub
positively saying its package seals no repository password (legacy hash-less
escrow). Nothing to compare, and offering would put a permanent screen in
front of every legacy box.
The write damper: CheckedAt refreshes on every ack carrying a hash, but a save is
skipped when both the hash and the UTC day are unchanged, so an idle box does not
rewrite settings.json every fifteen minutes. It records WHEN WE LAST HEARD, not
when it last changed - the R-100 distinction.
Tests: Scenario C (a differing key offers, with both proxies asserted false first),
Scenario D (a matching key offers nothing), fact 1 still required, shape (a) still
works, and both SEC 7.2 halves.
RED-PROOFS, each with the mutation confirmed present in the file first:
D) hubHash != localHash conjunct dropped -> Scenario D FAILS (a healthy box
offered recovery forever); Scenario C still passes
WIRING) RecordEscrowKeyHash removed from the EscrowAutoConfirmer literal in
main.go -> TestMainWiresRecordEscrowKeyHash FAILS. This is the ships-inert
shape: unwired, everything compiles, every test in the package passes, the
auto-confirm still works, and shape (c) reads an empty hash forever.
Green: go build, go vet, go test ./... all pass.
|
||
|
|
763de3a025 |
R-241 part 1: the box does not mint a repository key over a sealed package
THE DEFECT. WriteOffboxSecrets auto-generated on ONE input - does the file
exist. Its two neighbours in the same file, OffsiteRecoveryOffer and
needsOffsiteCredential, both consult GetHubEscrowIdentityPresent(). The same
fact was available on three paths and used on two.
Measured on the final walk: a rebuilt box's credential self-heal reached here
at 03:18:06Z and minted 9b4a9a9d over a hub package sealing 30ef574f. The
recovery screen then correctly reported nothing recoverable under the key the
box held. The screen was honest; the minting was not. And the flag was not
merely available at that moment - it was the PRECONDITION of the chain that
reached this function, logged at 02:48:03Z, six ticks earlier.
THE GUARD IS A CONJUNCTION, deliberately: a package held AND no key present.
A box the hub holds nothing for mints exactly as before.
The refusal is a HOLDING state, not a failure. ApplyOffsiteTarget catches the
sentinel and still writes the transport (ssh key, known_hosts, coordinates),
so the recovery screen can bring the tier up the instant the escrowed key is
placed (R-219). Returning the error instead would leave needsOffsiteCredential
true forever and the hub re-staging a consumed credential on every cycle.
New declared state offsite.state=awaiting_recovery_key, shown INERT to every
existing hub reader from their code rather than assumed: offsiteheal acts on
exactly one string; isStale needs Enabled && escrowed and this carries
Enabled=false; the delivery checker skips the applied shape; an unknown state
string is ignored by encoding/json. So NO hub change is needed for this part.
OffboxAwaitingRecoveryKey is DERIVED, not stored - the operator's ruling that
the state should be fixed rather than remembered, applied to this field too.
t.Enabled is load-bearing in that predicate and was MISSING in the first
draft. The existing TestOffsiteDeclare_DisabledTargetIsNotStranded caught it,
not review: a customer who switched off-site off is not awaiting anything.
Now pinned from the new predicate's own side as well.
Tests: Scenario A (no key written; transport still written; apply holds and
stages nothing), Scenario B (first-time box still mints), idempotency, the
nil-settings fail-safe, and the Scenario E carve-out.
RED-PROOFS, each with the mutation confirmed present in the file first:
A) guard block deleted -> both Scenario A tests FAIL with
"R-241 REGRESSION: apply minted a repository password over the sealed
package"; Scenario B still passes (the mutation is specific)
B) guard over-widened (hub-package conjunct dropped) -> Scenario B FAILS
with a first-time box unable to start; Scenario A still passes
Green: go build, go vet, go test ./... all pass; controller_gates all OK.
|
||
|
|
c6b69d888e |
v0.205.0 — a run that skipped an app the customer selected is not successful (R-234)
gates / gates (push) Successful in 21s
THE VERDICT. The R-203 block already said "a warning beside a success is read as a success" and applied it to ONE of the two shapes it describes: an app missing a declared mandatory FOLDER made the run incomplete, while an app skipped ENTIRELY still reported ok. Both do now. Which skips count, decided by measurement: selected+deployed with no recovery unit YES; selected but NOT deployed no (named, with what to do — a box left amber by an app somebody removed is a status nobody reads); disconnected/decommissioned drive no (own signal); nothing selected no. LastSuccess and SnapshotCount still record what WAS captured. THE FILED MECHANISM WAS NOT THE MEASURED CAUSE, and saying so is the point. §3 stated that toggling an app on leaves it without a bundle so the first run skips it. Measured on demo-hp: the run's own pre-dump phase calls captureAllRecoveryUnits for every DEPLOYED stack, through admitApp, before the push — a unit moved aside was RECREATED and the run reported ok. That state does not survive a run. What actually produced the 2026-08-06 sequence: the manual run was dropped by the single-flight while an earlier run was still going. runOffboxBackup returned nil, the handler had already answered "A tavoli mentes elindult", and the card then showed the PREVIOUS run's green verdict — read as covering the app just selected. The decision is now taken synchronously in the handler and a dropped request says so. The nightly path still returns nil on purpose: nobody asked, and it retries. §7.3 measured before deciding: CaptureRecoveryUnit writes a few KB of compose + manifest, only ENUMERATES dumps rather than creating them, is idempotent and does NOT stop the app — and already runs inside the off-site run. So there is no wait to remove for a deployed app and NOTHING was built. 28 packages ok, 9/9 gates. Four red-proofs, each asserted to have applied. Fixture note: the shared provider's ListDeployedStacks returned nil, so Scenario A first passed for the wrong reason; fixed with an opt-in deployed set that defaults to nil. |
||
|
|
53e9bf0224 |
v0.204.0 — the restore list is keyed on the store (R-237); the size gate stops refusing in silence (R-238)
gates / gates (push) Successful in 26s
R-237: /backups/restore listed apps that are CURRENTLY DEPLOYED and CURRENTLY TOGGLED ON for future off-site backups. A rebuilt box has neither, so a household that had just lost everything was shown nothing to restore while the repository held their snapshots — measured live on the R-201 re-walk. To restore an app you had to select it, to select it you had to have installed it, and to know what to install you had to see the backup you could not see. The store is now the source of the list (offsite_restore_list.go), built on the existing R-193 OffsiteInventoryList. Installed-ness became a property OF a row, never a filter on it. Every case is answered rather than hidden: a snapshot for an app that is not installed is offered and says it will reinstall first; an installed app with no snapshot is shown as having nothing; an unreadable store renders as UNKNOWN (R-225's rule, one screen over) AND keeps the action, because "we could not look" is not "there is nothing"; no-target is its own state. The felhom-offbox and _shares marker tags are excluded from the app list. R-238 classified as a HARNESS ARTIFACT: mode=full without confirm=1 is step 1 of a deliberate two-step — it starts no job by design and redirects carrying &full_prep=<app>, which deriveWizardStep requires to reveal the commit. A driver that did not carry it forward landed back on the intent step. The operator's browser run completed the same restore. The wizard's precedence rules were NOT re-keyed: a stale ?full_prep= must never resurrect a commit button mid-restore. The residue WAS real and is fixed: neither branch of that step wrote anything to the log, so a refusal — including by the headroom gate — left no trace on the box. Both branches now log, and so does the concurrent-op refusal. resolveWizardApp is removed: it was dead once the gate moved, and its test pinned the defect's behaviour (an untoggled app refused), which would have read as policy. 28 packages ok, 9/9 gates OK. Three red-proofs, each asserted to have applied. |
||
|
|
4d349d1106 |
REPORT + CONTEXT for v0.203.0: the retry shape, the marker answer, R-220's shape
gates / gates (push) Successful in 10s
Records the decisions rather than only the code: - POLL not ACK, decided on Scenario B against the ACTUAL promises — the no-target message gives no deadline and the card says 'within a day', so a 5-minute tick is inside both and no text needed changing. If either promise tightens to minutes, go ACK-driven. - The marker question: applied_marker lives in the guest's DataDir, which a rebuild destroys, so it cannot suppress a legitimate re-run. Left alone. - R-220 candidate (b), corroborated rather than a wider prefix, reading /proc/mounts because the lsblk args are pinned in sudoers. Live: Scenario C proven on demo-hp WITH a positive control — the job ran once and logged nothing. A first reading counted 2 lines that turned out to be the start-up reconcile, not the retry; the instrument was corrected before the conclusion. Scenarios A and E are deliberately NOT live-proven here: both need a rebuilt box, and that state arises naturally in Part 4. |
||
|
|
9dc26459ea |
v0.203.0: the box collects what the hub staged for it (R-218 consume half) + R-220's message
gates / gates (push) Successful in 10s
R-218's declaration half shipped in v0.201.0 and works. Its consume half never existed. Reconcile ran exactly twice per process — at start-up and when the recovery screen drives it — and BOTH fire before the hub has anything staged, because the hub stages in RESPONSE to the declaration those runs precede. Measured on the R-201 re-walk: unlock reconcile 11:43:07, hub staged 11:44:57 saying 'next cycle', a full report cycle ran 11:55:46, still unconsumed at 12:06. A guest command line applied it in 18 seconds — everything correct except the trigger. Bridge.RetryIfDeclared re-runs the SAME reconcile on a 5-minute tick, driven from the box's own published declaration (OffboxReportStatus().State) — the very statement the hub acts on, so the two cannot disagree. Poll, not an ACK flag, decided on the promise: the no-target message says 'amint megvannak' (no deadline) and the card says 'within a day'. Five minutes is inside both by a wide margin and needs no hub change. It stops by construction — a healthy box does no work and logs nothing — and the settle gate is deliberately kept via ReconcileWhenSettled. The marker was investigated and left alone: applied_marker lives in the guest's DataDir, which a rebuild destroys, so it cannot suppress a legitimate re-run. R-220's customer half: the refusal no longer tells the customer to choose from a list that may be empty. It names the rebuild, points at the Meghajtók page, and promises no outcome. Red-proofs: remove the retry -> credential uncollected (the dead end reproduced); drop the stop condition -> a healthy box hammers the hub; call Reconcile instead of ReconcileWhenSettled -> settle gate bypassed; restore the old sentence -> the impossible action returns. 28 packages ok, vet clean, all controller gates OK. |
||
|
|
66d80efb9f |
docs: R-168 is CLOSED — the "CI is still owed" sentence was stale (R-229 part 2)
gates / gates (push) Successful in 19s
Corrected in all four instruction files across all four repos. Found while confirming this session own push by run ID, which is precisely the check that catches it. In felhom-agent/CLAUDE.md the sentence contradicted the same file release section, which already said R-168 mails the failure -- a contradiction inside one instruction file, the exact class the R-229 work exists to find. REPORT.md deliberately NOT overwritten in the sibling repos: a one-line docs correction must not destroy the record of their last real implementation. |
||
|
|
7db42c5fec |
docs: CLAUDE.md becomes a core plus path-scoped rules (R-229)
gates / gates (push) Successful in 12s
215 lines -> 110 (92 effective; block-level HTML comments are stripped before injection and never reach the model, verified empirically on Claude Code 2.1.222 with a control and a treatment run). Four new .claude/rules/*.md, each with a paths: glob list so it loads only when a matching file is read: gates, ui-hungarian, backup-paths, agent-coupling. The ## Layout tree was deleted as derivable; REUSE.md already owns the per-package seams its annotations stood in for. The host/access table was deleted in favour of a pointer to documentation/operations/nodes.md -- it carried three defects at once: demo-felhom given as the LAN fallback address as if it were the route, a pinned "agent 0.93.0" against the project's own no-versions-in-docs rule, and the claim that no drill VM was provisioned on demo-hp. Measured live: qm list shows VM 300 drill-r50. felhom-agent/CLAUDE.md was right; this file was wrong. Kept verbatim: the seven session-critical invariants, the F9 live-validation fence, the end-of-session checklist. controller_gates.py registers the shared instructions gate (felhom.eu/scripts/, never copied here; an absent sibling clone FAILS). Docs only -- no Go, no version bump, no image, no deploy. Ledger: felhom.eu/documentation/audits/LEDGER-instruction-trim-2026-08-06.md Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JJc8sAGRWmavP3rMtdpkr2 |
||
|
|
a62bb3874b |
REPORT + CONTEXT: v0.202.0's rule, its live proof, and what was NOT verified live
gates / gates (push) Successful in 9s
CONTEXT gains the rule so it outlives the bug: on the unlock path the customer is blamed only after a real attempt REFUSED their code; every other outcome, including an unclassifiable one, says something else. Plus the two things that must not be 'fixed' into it — elapsed time is never a classifier, and the error TEXT is never read (when the distinction was not a value, the agent was changed to provide one). REPORT states the split honestly: the AGENT half is proven live on the venue (400 -> 502 -> 400, same wrong code, only the hub's reachability changed), while the controller's message selection rests on handler tests and red-proofs, because /recovery correctly redirects since F7 set the old data aside and restoring that state is the reconfiguration §11 forbids. Also records that the correct codes were shredded by the previous session, so the live re-run used a WRONG code — which makes the test harder, not weaker. Two venue changes stated because they were not asked for, both restorations: a fresh dashboard password (the previous session shredded it, leaving the box impossible to log into) set through the supported --print-reset-code escape hatch, and one normal off-site run to populate stats_known. |
||
|
|
7534ea203d |
CHANGELOG: controller v0.202.0 (MinAgent 0.126.0)
gates / gates (push) Successful in 10s
Declares the coupling: an agent below 0.126.0 answers 400 for both a fetch failure and a wrong code, so FeatureRecoveryFailureClass withholds the refusal reading and the 400 degrades to the neutral message. The gate blocks nothing — it only decides whether the customer may be told to check their typing. |
||
|
|
c7446f2d6a |
R-225/R-227/R-228 Parts 2-4: unknown is not zero, the gateway speaks Hungarian, the set-aside is visible
R-225 — an unread store said '0 pillanatkép / 0 / 50 GB' above a card stating it held backups under another key. An SFTP listing found snapshot f3d9cd67 and 12 535 KB really there; snapshot_count and repo_size_bytes were simply ABSENT and the zero value spoke for them. StatsKnown is now NAMED, for the same reason OffsiteInventory.Empty is: zero is what an unread store and an empty one both look like, and on the wire 'absent' and '0' are the same bytes. The fill bar renders only when the fill is known — a 0%-wide bar is a picture of emptiness, and a picture is a claim. A measured zero still says zero. R-227 — WHICH LAYER ANSWERS: traefik, and this repo generates its config. But traefik v3 serves no static files, so a branded proxy page needs a new always-up container for every 502 on the box — out of proportion, and scoped in the report rather than built. Shipped instead: the unlock posts via fetch and answers a gateway failure in Hungarian without leaving the page. Progressive enhancement — with no JS the plain POST is unchanged and still shows the proxy's error, which the report says plainly rather than implying otherwise. R-228 — the set-aside history was recorded in orphaned_renamed_to and read by nobody: a census found zero references in any template or handler, while 12 535 KB sat at that path. It is surfaced as two facts and stops. It does NOT promise the history can be reopened, because it cannot be by anyone today (R-199's inventory is unbuilt) — and the set-aside CONFIRMATION copy was corrected for the same reason: 'a helyreállítási kód nélkül többé nem lesznek megnyithatók' implied that WITH the code they could be. The field's own comment called it 'recovery-code-recoverable', which was the same over-promise in the code. Tests: scenarios F, G, H as render tests per branch of each gate. Red-proofs, each demonstrated failing then restored: remove the StatsKnown guards (F, 'R-225 RETURNED: an unread store reports a snapshot COUNT of zero'), delete the set-aside block (H). The F assertion on the fill bar is scoped to the bar's own container — a bare width:0% search matched unrelated elements and would have passed for the wrong reason. 28 packages ok, vet clean, all controller gates OK (the emoji gate caught a warning sign in a template comment). |
||
|
|
1e759a16ec |
R-224/R-226 Part 1: why the unlock failed decides what we say
The failure branch was a two-way choice — superseded? M4 : M1 — and BOTH are statements about the customer's code. rerr was never inspected, so a hub that refused, an agent that was stopped and a genuinely mistyped code all produced the same accusation. Measured live 2026-08-05 with a CORRECT current code: hub firewalled off 0.0556s, agent stopped 0.0299s, against ~1.0s for a real unseal. Five classes, from the VALUE and never the text: hub-unreachable 502/503 from the agent — the code was NOT used agent-unreachable no agent verdict at all (transport) — NOT used no-bundle 404 bundle-too-old 409 asked-and-refused 400 — the ONLY class that may mention typing unknown everything else -> NEUTRAL, the safe default agentapi.RecoveryRefusal carries the status as a value (refusalError flattened it into a sentence, and a sentence is not something a caller can branch on). THE OLD-AGENT CASE IS WHY THIS NEEDS A COUPLING. Agent < 0.126.0 answers 400 for both a fetch failure and a wrong code, so a 400 from one cannot be read as a refusal. FeatureRecoveryFailureClass (MinAgent 0.126.0) withholds that reading and the 400 degrades to neutral. The gate BLOCKS NOTHING — it only decides whether the customer may be told to check their typing. R-226: the superseded message now names BOTH possibilities and restores the ten-words prompt. The two are indistinguishable at the engine; the honest message says so. It still does not promise the earlier package can be opened. Elapsed time is logged (it is what diagnosed this) and is NEVER a classifier. Tests: scenarios A-E at the HANDLER + the classifier table asserting the same sentence under two statuses classifies two ways. Red-proofs, each demonstrated failing then restored: delete the 502 case (A), remove the mistype clause (C), default to the accusation (D), route an instant transport failure to the typing message (E). Two existing tests encoded the defect and were corrected, not deleted: the web fake returned a BARE error for 'wrong code' (which is the shape of a failure we cannot classify), and R-222's test forbade any mention of typing on a superseded box — half of which R-226 deliberately reverses. 28 packages ok, vet clean, all controller gates OK. |
||
|
|
05cf352a2f |
docs: CAMPAIGN-11 fix pass — REPORT + CONTEXT (controller v0.201.0)
gates / gates (push) Successful in 8s
|
||
|
|
a3499d1807 |
v0.201.0 — a correct recovery code is never called wrong again (CAMPAIGN-11) — MinAgent 0.125.0
gates / gates (push) Successful in 9s
R-216: the offsite key recovery is a coupled feature and now says so. featureProbes +
featureMinAgent 0.125.0 + a Supports gate at the unlock entry point, FAILING CLOSED — an
agent that cannot answer is named as such instead of the customer's code being blamed.
Measured live: a 404 from agent 0.120.0 came back as "we did not accept your recovery
code, check that all ten words", in 0.134 s, against a perfect code.
R-218: delete the repo-password short-circuit in needsOffsiteCredential. The declaration
stops when the TIER WORKS, not when a key exists — installing a key is the recovery
screen's whole job, so succeeding at recovery was switching off the mechanism that would
have delivered the coordinates to use it.
R-219: the unlock finishes the job — place the key, bring the tier up, then list. Without
it the promised listing could never render on the shape the screen exists for.
R-217: an unreadable store no longer claims to have opened with unattributable content
(the OffsiteInventory{} zero value). Opened / empty / unreadable are three states.
R-222: a code that is right about a RETAINED earlier package is named, not blamed. States
what the hub knows and promises nothing — no read path exists.
R-215: GET /recovery is gated on the same predicate as the interception.
Five red-proofs, each demonstrated failing and restored.
|
||
|
|
a315d623b8 |
docs: R-193 CLOSED — CONTEXT + REPORT (controller v0.200.0)
gates / gates (push) Successful in 10s
|
||
|
|
62b85ecf13 |
CHANGELOG: controller v0.200.0 (R-193, the recovery screen)
gates / gates (push) Successful in 9s
|
||
|
|
636c51e542 |
R-193: the recovery screen — unlocking, and only unlocking (v0.200.0)
A customer whose machine was rebuilt had everything needed to get their data back and no way to find out: the only route was a command line. This is the screen that closes that. IT UNLOCKS, AND ONLY UNLOCKS (operator ruling). It explains, takes the recovery code, opens the repository and shows what is in there — apps, dates, sizes. It restores nothing: restore is already per-app and lives in the backups area, and a screen that unlocks and then offers to overwrite is two decisions wearing one button. ONE CORE, TWO CALLERS. RecoverInstallCore is split out of RecoverAndInstall; the CLI wrapper keeps its exit codes and printed lines byte-identical, and the handler drives the same function. Two implementations of the one operation that can permanently lose a customer's data would drift, and only one would be tested. Asserted from source on both sides by AST. THREE WAYS OUT, none a dismiss button: recover; 'most nem' (the full page stops interrupting, the backups-area entry point stays PERMANENTLY, bound to the offer and never to the postpone flag); and 'I do not want the old data' — confirmed TWICE and reaching the SHIPPED move-aside, which sets aside and never deletes. THE CODE IS HANDLED NO MORE LOOSELY THAN ON THE COMMAND LINE: POST body only, never logged, never persisted, never echoed, cleared on every path, no-store, autocomplete off. No lockout — the code is a ten-word phrase, and locking a customer out of their own data for a typo is worse than anything it prevents. TWO DEFECTS THE TESTS CAUGHT, both fixed: an UNCLAIMED (legacy-open) box would have been shown the page, because RequireAuth passes such a box through; and the inventory nil-dereferenced when no off-site target was configured, which is exactly the pristine rebuilt shape. |
||
|
|
be3c5fa7f6 |
docs: R-204 item 4 (box half) — CONTEXT + REPORT (controller v0.199.0)
gates / gates (push) Successful in 10s
|
||
|
|
992803c10b |
CHANGELOG: controller v0.199.0 (R-204 item 4, box half)
gates / gates (push) Successful in 10s
|
||
|
|
a91f055960 |
pre-push: refuse a push from a clone outside the felhom workspace (R-204 rider)
Identical to the assertion added in felhom-agent 0404f60 and app-catalog-felhom.eu ee2c810. See those commits for the reasoning. |
||
|
|
1214bae0a2 |
R-204 item 4 (box half): a rebuilt box DECLARES that it needs a credential (v0.199.0)
An absent off-site object has four meanings — never configured, mid-restart, a transient config read failure, and rebuilt-and-stranded — and the hub cannot tell them apart. The box can, from two local facts it holds with certainty, so it says so instead of leaving the hub to deduce it from a silence (operator ruling). The ACK's identity_blob_present is now recorded on EVERY ACK, before the gates that used to discard it: on a box with no off-site target the auto-confirm returns immediately, which is exactly a rebuilt box, so the one fact distinguishing it from a box that never had off-site backups was thrown away every cycle. The declaration needs BOTH halves — a fresh data area AND a hub-held recovery package. Freshness alone is a box that never had off-site backups; dropping that condition makes the whole fleet ask for credentials, which is what the Scenario B test exists to catch. The object carries enabled:false and zero sizes, which is what makes it inert to the hub's existing fill and staleness checkers and to a pre-upgrade hub. A configured box's JSON is byte-identical to v0.198.0's. |
||
|
|
68f195676b |
docs: R-204 items 1 & 3 — CONTEXT, REPORT, README (controller v0.198.0)
gates / gates (push) Successful in 9s
|
||
|
|
33fcc502e4 |
CHANGELOG: controller v0.198.0 (R-204 items 1 and 3)
gates / gates (push) Successful in 10s
|
||
|
|
2e936f43bf |
R-204 item 3: a restore says what it restored, and what it did not (v0.198.0)
mode=unit restores the recovery unit — the app's definition, configuration and database dumps — and NOT the customer's own files: RestoreOffboxScratch passes --include <unit path> and the userdata in the same snapshot is excluded by it. The outcome was one sentence for both modes and named neither scope, so on the last step of a disaster recovery the customer was told the app had been restored after the thing they were looking for had not been. restoreScratchOutcomeMsg states what came back, what did not, and the next step that gets it. The wizard's intent card states its scope before the choice. The full-restore size gate is untouched and pinned as unchanged; the default stays unit, since all three wizard forms set mode explicitly. |
||
|
|
73b6dbc27d |
R-204 item 1: a freshly minted reset code works without a restart (v0.198.0)
--print-reset-code runs as a separate process and persists the new code; the running server's cache was never told, so the code the customer was told to type was refused until the controller restarted. Nothing said so — during the 2026-08-04 drill that cost two attempts with an operator present. effectiveClaimCode now reads through to the persisted state before applying the settings-vs-config precedence, which is itself unchanged. Read-through, not a TTL: a TTL would leave a window in which a superseded code still works, which is worse than the bug. Fails closed on an unreadable state; an absent file is not an error. |
||
|
|
f4796e0d00 |
docs: R-203 contract + report (controller v0.197.0, proven live)
gates / gates (push) Successful in 10s
|
||
|
|
58c703bd44 |
R-203 Part 2: a run that missed a MANDATORY directory is not a successful run (v0.197.0)
gates / gates (push) Successful in 8s
The gap was already detected and warned about, in Hungarian, naming the app and the folders -- that warning is what stopped the R-201 drill. The defect was that the run still reported `ok` beside it, and a warning standing beside a success is read as a success. last_status gains "incomplete": minted, because "ok" | "error" | "running" had nothing meaning "it ran, and this app is not fully protected". NOT "error" -- the rest of the run worked and what was captured is real, so SnapshotCount and the LastSuccess anchor still record it. Half a backup is not no backup. The gaps are now recorded STRUCTURALLY (offboxRunResult.mandatoryGaps), not only as prose, so the verdict has something to act on. It reaches the operator through the EXISTING per-run digest (backup_run_failures) rather than a new event type -- a new type is a two-repo change and the hub drops anything outside allowedEventTypes. The stat-filter gains the ClassMandatory check Tier 2 already had. It is a NO-OP today (TierOffsite admits mandatory only), so no customer-visible warning disappears -- demonstrated by widening the tier filter alone and watching the check hold the line. ANTICIPATED: calibre-web on demo-hp has exactly this gap, so its off-site status becomes incomplete the moment this ships. That is correct and is the point. Red-proofs: my first Scenario-C proof PASSED because the test only reached offboxCaptureSet while the mutation lives in runOffboxInternal -- a mutation the test cannot observe is not a red-proof, and the fix was the test. The run-level test now fails under both mutations (unreachable gap recording; unconditional ok). |
||
|
|
a96c3d9473 |
R-203: the export-mount resolver takes the namespace root too (its own commit)
gates / gates (push) Successful in 9s
ExportDataMounts lives in delete.go, which reads as a destructive path. IT IS NOT: its single production caller is the .fab export adapter, and nothing deletes based on its result. The delete path's own guard, ProtectedHDDPaths, is layout-agnostic by construction -- it protects BOTH <hdd>/... and <hdd>/felhom-data/... -- so deletion was never affected by the namespace-root defect. That scope note is now in the function's doc comment, because the file placement will mislead the next reader exactly as it misled the spec for this change. Separated into its own commit anyway, so a change to a function whose filename says "delete" is reviewable on its own. An empty nsRoot falls back to hddPath -- the pre-R-203 shape -- so any caller not yet updated keeps working on enrolled drives. Tests cover both drive kinds and assert the NEGATIVE: no emitted path lies outside the app's own data roots. Red-proof: leaving the site bare fails the system-drive row, emitting /mnt/sys_drive/userdata where the canonical root is /mnt/sys_drive/felhom-data/userdata. |
||
|
|
73efb091d9 |
R-203: the app and its backup look in the same directory — one resolver, every caller
gates / gates (push) Successful in 9s
appbackup's path helpers take a NAMESPACE ROOT. Five call sites passed a bare DRIVE path.
On an enrolled drive the two coincide, so nothing showed; on the system-data fallback they
differ by exactly the felhom-data segment, and the app then bound a directory the off-site
capture set never looked at -- while the run reported ok. Measured live on demo-hp: the app
wrote to /mnt/sys_drive/userdata/media/books, the capture set looked for
/mnt/sys_drive/felhom-data/userdata/media/books.
THE RULE NOW HAS ONE EXPRESSION. appbackup.NamespaceRootFor / IsEnrolledDrive encode the
drive-kind comparison; backup.Manager.namespaceRoot and stacks.Manager.inGuest delegate to
it. There were already TWO copies and they differed -- the backup package's compared without
filepath.Clean, the stacks package's with it, so a trailing slash from config would have
flipped the mode in one and not the other.
Sites routed through it:
- stacks/deploy.go withPathVars -> ${USERDATA_PATH} (the live defect)
- appexport/fabplan.go + export.go (via a new provider method)
- web/handlers.go FileBrowser mounts (latent: the system drive is
deliberately never a registered StoragePath, so this is the identity today)
ComputeFabBuckets now receives the namespace root, which is what ComputeCaptureSet has always
received -- so the export's classified paths and the backup's capture set describe the same
directories by construction instead of by coincidence.
Tests are table-driven over BOTH drive kinds, because this survived by being invisible on the
kind that already worked. Red-proofs observed: restoring the bare-path call fails the
system-drive row with the two paths differing by /felhom-data; inverting the drive-kind
comparison fails every enrolled row.
|
||
|
|
532f5712a8 |
docs: R-200 Part 0 shipped; R-203 recorded (mandatory userdata dir missing from the offsite snapshot while the run says ok)
gates / gates (push) Successful in 9s
|
||
|
|
1b1366bb6e |
controller v0.196.0: the recovered key installs itself (R-200 plumbing half) -- MinAgent 0.125.0
gates / gates (push) Successful in 8s
--recover-offsite-install is the sibling of --recover-offsite-check: same fetch/unseal path through the agent, same STDIN discipline for R, but it PLACES the recovered repository password via InjectOffboxPassword so a rebuilt box reopens the history it inherited. Doing this by hand would put the offsite DATA key through a terminal, a clipboard and shell history. In-process the value goes agent -> this process -> the 0600 file and is rendered nowhere. The confirmation is a SECOND invocation: without --confirm-install it prints both hashes and writes nothing, so the operator sees the comparison before any write is possible. Three outcomes, named distinctly: installed (no local password -- the rebuilt-box shape), unchanged (identical key already present, nothing written), refused (a DIFFERENT key present; installing would clobber the key the current repository is encrypted under, and no force option is offered). Exit 2 for the refusal, distinct from 1 for a failed step. Red-proof: removing the confirmation gate makes the dry run write, failing the test. The R-persistence test carries a positive control -- a planted copy is found, then removed and not found -- because an absence check is worth only what its sensitivity is. |
||
|
|
bdab80c933 |
docs: R-200 diagnostic — CONTEXT + REPORT (proven live on demo-felhom)
gates / gates (push) Successful in 10s
|
||
|
|
9640e51321 |
controller v0.195.0: prove the offsite key comes back (R-200 plumbing half) -- MinAgent 0.125.0
gates / gates (push) Successful in 10s
--recover-offsite-check is a docker exec diagnostic in the shape of --print-reset-code: it reads the customer's recovery code from STDIN, asks the agent to fetch this host's sealed bundle and open it, and reports whether the recovered key matches the one on disk BY SHA256. Two hashes and a verdict; never a password, never R, never a blob. R comes from stdin and not a flag because a flag value is visible in ps, in shell history, in a container's command line and in any transcript of the session that ran it. IT COMPARES; IT DOES NOT INSTALL. The recovered password is never written to offbox/repo_password -- installing changes a live box on a path nobody has walked, and that link is next session's, with the drill around it. A test asserts the data dir is byte-unchanged after a check; its red-proof (adding the install call) fails it. Exit codes: 0 match, 2 clean MISMATCH, 1 a step failed -- "it failed" and "it worked and disagreed" must never share a status. A box with no local password reports distinctly: that is the rebuilt-box shape, where the next step is to install rather than compare. Nothing customer-reachable ships here: no card, no form, no preview. |
||
|
|
0887fd676d |
REPORT: R-182 — the run digest, the live proof, and the red-proof that did not fail first time
gates / gates (push) Successful in 9s
|
||
|
|
88897a224e |
v0.194.0 — one operator email per backup run, and nothing dropped without a trace (R-182)
gates / gates (push) Successful in 8s
MEASURED, not supposed. On 2026-08-03 nine per-app recovery_unit_capture_failed events reached the hub and TWO operator emails went out. The hub's operator cooldown key is customerID:eventType(+tier) and that event carries `app` but no `tier`, so the key held no app identifier: the first refused app took the hour's slot and every other app's failure was discarded BEFORE anything was written down, leaving no row on any channel. The obvious fix — put `app` in the key — was ruled against: on a full disk it produces one email per app, the volume problem wearing the correctness problem's clothes. internal/backup/runsummary.go: a per-run collector with exactly admissionSet's lifetime, fed by all three write legs, emitting backup_run_failures ONCE at the end and only when something failed. A clean run emits nothing. The per-app event stays and becomes the RECORD — the hub routes it record-only, stored and logged every time, never competing for an email slot. The record and the notification are now different things. Deliberate skips (disconnected, decommissioned) are excluded: they have their own alert, and a nightly email about an unplugged drive is one the operator learns to ignore. A manual run always reports: the digest carries a unique run_id the cooldown cannot collapse. Someone pressing the button is actively trying to get a backup. THE PERIODIC SWEEP GETS A DIGEST TOO. With the per-app event now record-only, a capture failure found between runs would be recorded and never notified — a new silence introduced while closing one. That path emits a digest with NO run_id, so the ordinary 1-hour cooldown caps it exactly as before while the mail now lists every failing app instead of whichever was first. A refusal is recorded ONCE, where the verdict is taken, not at the three legs that consult it — R-181's contract is one verdict per app per run. Noting it per leg listed one refused app three times and produced "2 of 1 apps failed". Found by the digest's own test, not in review. Silence is safe because the hub's deadline check raises expected_backup_missed from report freshness, independently of any mail this box sends (monitor/deadline.go:396,417). Confirmed, not assumed. 7 new tests, 4 red-proofs. The main.go seam walk did NOT fail on its first attempt — the AST test walked the backup package and not main.go; the test was fixed and the mutation re-run rather than the pass recorded. |
||
|
|
db0d4b129d |
REPORT: R-181 — the reserve, the live proof, the du measurement and the teardown
gates / gates (push) Successful in 9s
|
||
|
|
6c43bf6156 |
v0.193.1 — the refusal's size estimate is rendered in bytes, not "0.00 GiB" (R-181 follow-on)
gates / gates (push) Successful in 9s
Found by v0.193.0's own live proof run. The estimate was printed fixed to two decimal GiB, so every app under ~10 MB rendered as "estimated 0.00 GiB write" — which reads as "no estimate was available" and is the opposite of what happened. Observed live on demo-hp 08:59:46: opengist's real 178 KB estimate printed as 0.00 GiB. Shipped in the same session because it is the same defect class R-181 is about: a message an operator cannot rely on is worse than no message. The arithmetic is unchanged and still in GiB — the reserve's own unit, so the comparison against FloorFreeGiB reads directly. Only the rendering moved to humanizeBytes. estimatedWriteGiB -> estimatedWriteBytes, with the GiB conversion done once at the point of comparison. |
||
|
|
fef07c3923 |
v0.193.0 — the reserve guards the write that fills the disk, and its promise is true (R-181)
gates / gates (push) Successful in 9s
B2's capture floor (v0.192.0) was consulted in exactly ONE place — captureAllRecoveryUnits, which writes a few KB. The two legs that write the BULK into the same backups/primary/<app> tree, the DB dump and the volume dump, ran FIRST and unguarded. Measured live on demo-hp 2026-08-03 06:40:03: opengist's volume dump wrote 2.0 GB with no check, free fell to 1.0 GB, and the floor then refused the cheap write it had already lost the argument to. Its refusal message claimed "the previous unit is untouched" — measured false: that app's tar had gone 182,272 B -> 2,147,666,432 B under a stale manifest. Sixth entry in CLAUDE.md's table of shipped guarantees the code did not provide. Fix: ONE admission verdict per app per run (internal/backup/admission.go), taken before that app's FIRST write and covering all three legs — they write under one per-app root, which is why one verdict can honestly cover them. - Lazy, at the app's first write, NOT once at run start: app A's dump can put app B under the reserve, so a run-start verdict reads a disk that no longer exists. - Remembered for the run, never re-decided between an app's own legs — that is the split this closes. Reset per run. - Placed ahead of DumpAppVolumesSafe, which stops the stack as its first act, so a refused app is never bounced. After the volume-less check, which has no write. - Exactly one operator alert per refused app per run. - Leg order unchanged: volume dumps still precede the capture. The floor is now SIZE-AWARE: it asks whether THIS app's write would cross the reserve, not only whether the filesystem is already below it — which is how an app was admitted at 96% and then allowed to write 2 GB. Estimate = the app's previous .sql + .tar on disk. No history -> headroom-only, deliberately, and the alert says so. A container-based du per volume was MEASURED and rejected: 66 timed runs on demo-hp guest 9201, median ~355 ms/volume (341-404) on volumes holding tens of KB — container start-up, not the walk. Decisive on top: docker run needs the writable layer, so it can fail under exactly the pressure the reserve handles. The message was NOT weakened; the behaviour was moved so the wording became true. It now also names which term bound. Every claim is checked against a sha256 fingerprint of the tree it describes, never against the log line. Still refuses and never deletes: nothing here is generational. 11 new tests through the production functions. The DB leg cannot run without Docker, so its gate is pinned by an AST walk of backup.go asserting admitApp precedes DumpOne (strings.Contains is insufficient — a commented-out call still contains the string). 4 red-proofs demonstrated failing then restored. |
||
|
|
4be6467b50 |
v0.192.0 — the capture floor replaces the bulkhead (R-165, decision B2)
gates / gates (push) Successful in 8s
Ships BEFORE the disk-layout merge it exists for, and is harmless on a box that never gets it. The mp1 partition was a BULKHEAD as well as a ceiling: it kept a runaway capture from filling the space the container runtime needs, because /var/lib/docker was a different filesystem. After the merge it is the same one, and a full Docker data-root is a stopped box. The floor sits in captureAllRecoveryUnits, checked BEFORE anything is written: below the reserve, that ONE app's capture is refused, its previous unit is left byte-identical, the R-158 alert fires with the space figures, and the loop continues. Two terms whichever binds first (97% used / 1 GiB free) in fillwatch's shape, deliberately BEYOND its critical band (95% / 2 GiB) so the customer is always warned before a refusal can happen — a floor that fires before its own warning is a silent failure wearing a threshold. Headroom, never unit size: a per-unit cap would be R-163 rebuilt inside one volume. Refuses, never deletes: nothing here is generational, so pruning could only destroy a different app's only local copy; pruneStalePrimaryDirs is an orphan sweep, not retention, and must not be repurposed. Tests 1184 -> 1191. One fixture strengthened mid-red-proof: the "old 20 G ceiling is gone" test sat at exactly 20 GB and survived a literal UsedGB > 20 cap — hollow. Now 120 GB, and the mutation fails it. |
||
|
|
d5be67b913 |
REPORT: CI run ids and conclusions (all three commits green)
gates / gates (push) Successful in 9s
|