1a1b32b3dde6261b93bef570f5c930563e4c13b3
1210 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1a1b32b3dd |
REPORT + R-437: the beta stopping line recorded, and the live trigger declined with its reason
gates / gates (push) Successful in 19s
REPORT.md carries the deployed sentence quoted from the RUNNING binary (kubectl cp + byte grep, both controls), the stopping line as it reads in all three places, the enumerated deferred set, the two provider questions, the register census, and the ArgoCD verification. R-437 filed: the register compression sweep is OWED and was deliberately not run here. Measured first — 12 of 181 rows / ~25 KB of 316 KB (about 7%) carry a closed leading verdict — so it buys little and touches everything, and it is the exact operation that misfiled seven rows in August (R-378; the seventh, R-87, sat wrong for nine days, R-405). The row carries the scope so it can be picked up cold. The live alarm trigger was NOT run and the report says so in its own section rather than substituting quietly: this alarm only fires on a real fall in a real customer's snapshot count, so firing it means either deleting real backups or POSTing a falsified report claiming demo-hp lost its own. That would write a fabricated point into a customer's report history, move its latch and baseline, and mail the operator a second alarm about a real box hours after the first one already confused him. Covered instead by the deployed-bytes proof plus three red-proofed tests driving saveReport -> Check -> notify. What remains unproven is named: that the dispatcher delivers THIS wording to a mailbox. Register 688 -> 700 lines; 182 rows; open-state 170. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM |
||
|
|
0f65f7a197 |
manifests: hub 0.111.0 -> 0.111.1 (R-434, the alarm's withdrawn promise)
gates / gates (push) Successful in 18s
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM |
||
|
|
db38f4c800 |
hub v0.111.1: the alarm stops promising a rescue that does not exist, and the arc is closed for beta
gates / gates (push) Successful in 17s
R-434 CLOSED — and the row's own "blocked on R-433" verdict was wrong, which is the point.
The fix is a DELETION, not a replacement: withdraw the promise instead of swapping it for a
new one, and the sentence is true under every possible answer to the provider questions, so
it never needs a second rewrite. A replacement would have been blocked; a withdrawal is not.
was: "...still hold the older copy, so this is recoverable file-by-file; it is NOT
confirmed data loss. Check whether a deletion ran on the box before restoring."
now: "...still hold the older copy. The route back out of them is not yet established,
so treat this as neither confirmed data loss nor confirmed recovery. Get in touch
before restoring anything, and check whether a deletion ran on the box."
It must not swing the other way either: "your backups are gone" is still usually false.
Clause (a) — the box cannot WRITE into the snapshot area — stands and is re-confirmed.
Tests: offsite_r434_test.go, three, all driving the production path so they assert the
sentence an operator RECEIVES. ASCII-only fragments, positive and negative controls.
RED-PROOF: restoring the v0.111.0 sentence failed all three, on every fragment, with the
offending sentence printed. TestR431_FiresOnAMassDeletion asserted "NOT confirmed data
loss" and caught this fix correctly; its wording fragment is REMOVED rather than updated,
so the wording keeps ONE home.
R-435 written into the detector's own documentation, no threshold changed: it sees a mass
deletion, not one app being wiped (69 across 9 apps -> ~35 needed, one tag is ~9, and
forget --prune groups by host,tags). Says explicitly not to lower the numbers.
THE STOPPING LINE, in all three places — register, 07 section 8 head, STATUS.md.
Deferred set ENUMERATED, not described: 07 section 8 rows 4, 8, 9, 10, 11 (+11b), 12,
each tagged [BETA-DEFERRED]. A number in the brief was wrong and is corrected in place:
six rows are DEFERRED, ELEVEN carry a blank RTO (4,5,8,9,10,11,11b,12,13,14,15); the other
five are blank for reasons that are not deferred work, and row 15 is an open DEFECT (R-104)
that the stopping line does NOT cover. NO STATUS MOVED — nothing was proven today.
Two provider questions drafted, not sent, no API called (11-D stands):
documentation/runbooks/provider-questions-2026-09-01.md, linked from R-95 and R-433, and
tracked by a dated DUE-CHECKS row (2026-09-15) — the 2026-07-27 check that sat unconfirmed
for 36 days is the scar that block exists for.
R-95, R-433 BLOCKED-ON-PROVIDER. R-95's one-day demotion on a clause that did not hold is
recorded; the proposal to rank it back near the top is stated and NOT acted on. R-430 marked
LATENT with its trigger: it becomes live the moment delete is withdrawn, so it is a
precondition on the R-95 build, not a follow-up. The stale ranking paragraph ("armed",
"zero snapshots") is corrected in place, order unchanged.
Register 621 -> 688 lines; 181 rows throughout; open-state 170 -> 169.
No controller or agent change. No golden owed, no floor change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
|
||
|
|
10c223bdfe |
DRILL R-95: the recovery route does not exist — stopped before the destructive phase
gates / gates (push) Successful in 19s
The drill was to delete demo-hp's off-site history and get it back out of a Storage
Box snapshot, filling row 10's blank RTO. Phase 1 found there is nothing to get it
back from: no snapshot is reachable from a sub-account BY ANY NAME.
Measured, read-only, no delete verb issued against any live store:
- 777,600 exact names in the vendor form YYYY-MM-DDTHH-MM-SS, nine full days at
second granularity, plus 126 alternative shapes -> ZERO hits.
- The control is what makes that mean anything: the identical 600-name batch shape
with one real path appended returned it, 6 of 6.
- Structural cause: /home (u629488-sub3) is st_dev 0,82; /.zfs/snapshot is st_dev
0,276; /home/.zfs does not exist. A snapshot under /.zfs/snapshot belongs to a
different dataset than the one holding felhom-repo.
- Three tools agree with controls in the same run: SFTP, the port-23 shell,
rsync --list-only.
So yesterday's re-scope splits: clause (a) "the box cannot write into the snapshot
area" STANDS and is re-confirmed; clause (b) "the rest is recoverable file by file"
is NOT SUPPORTED. STOPPED before Phase 2 on the operator's ruling — with no recovery
leg the deletion would have destroyed real history to buy only an alarm test that
could not fire at the specified size. Store verified untouched at 69 snapshots.
R-432 ANSWERED (negatively; its panel-read next step withdrawn as unnecessary).
R-433 no snapshot reachable by any name — decides R-95's remedy and its rank.
R-434 the drop alarm's text promises a file-by-file recovery that cannot be performed.
R-435 the drop detector is blind to a single-app deletion (>50% of 69 needed, ~9 given).
R-436 LEAD: the provider offers `rclone serve restic --stdio` and restic 0.14.0 speaks
`rclone:` (measured, controlled) — real prevention may need no new machine, IF
the vendor pins --append-only. Ask before building.
07 §8 row 10: text corrected, status NOT moved, RTO still blank.
No code, no version bump, no image, no golden. REPORT.md's only copy of the R-331
report preserved as REPORT-r331-backup-card.md before overwrite.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
|
||
|
|
17d92e71a1 |
REPORT + CONTEXT + STATUS: the record corrected, and the thing worth building shipped
gates / gates (push) Successful in 18s
Opens with Part 1's answer because everything reads differently after it: a customer's own account can reach the snapshot DOOR and is REFUSED writes to it, but sees the tree EMPTY. The write-refusal is the load-bearing sentence of the whole R-95 re-scope and it is now PROVEN rather than cited - the control write to the account home succeeded and was cleaned up, the write into /.zfs/snapshot returned `dest open ...: Failure`, and nothing was left behind. Identical on both boxes. storage-box-pool-1 IS u629488, so the emptiness is per-sub-account filtering rather than absence - which means recovery is an operator act in a browser today (R-432), and that decides whether R-95's remedy can ever be product-driven. STATUS carries two items for Viktor in plain words: read one snapshot name off the panel (two minutes, and it may make recovery product-reachable), and IGNORE the alarm mail he received today - the live firing was required to prove delivery and nothing was deleted. Five of my own mistakes are named, including the one that matters most: my first escalation-only test was HOLLOW and its red-proof PASSED. It re-swept the same report, so the baseline had already moved and the latch was never consulted. That is why red-proofs are run. |
||
|
|
65c82c4aa0 |
manifests: hub 0.110.0 -> 0.111.0 (R-431)
gates / gates (push) Successful in 16s
The manifest is the truth - a code push and an image build deploy NOTHING until this tag changes in git AND the app is synced. Auto-sync is OFF. |
||
|
|
30681764cb |
hub v0.111.0: notice a deletion within a day (R-431); correct R-429; re-scope R-95
gates / gates (push) Successful in 17s
THE RECORD WAS TELLING A WORSE STORY THAN THE TRUTH FOR TWO MONTHS, and my own probe is why.
R-429 CORRECTED. Yesterday's spike searched for a directory called `.snapshots`. The vendor documents
the path as /.zfs/snapshot. The probe's CONTROLS were sound and its SUBJECT was wrong, so "not found"
was true and meant nothing. Re-probed at the documented path on both boxes, with controls:
- /.zfs lists (shares, snapshot) from inside the jail;
- a write into /.zfs/snapshot is REFUSED - dest open ...: Failure - while the identical write to
the account home SUCCEEDS and was cleaned up.
That is the append-only property PROVEN rather than cited, and it is the sentence the whole re-scope
rests on. Seven daily snapshots are confirmed in the panel. The mitigation works. What remains true,
and was always the actual finding: the row claiming it had no R-number, its "confirm tomorrow" went
36 days unanswered, and the DUE-CHECKS block built for that class was empty. THE FINDING WAS NEVER THE
SNAPSHOTS - IT WAS THAT NOBODY COULD TELL.
R-95 RE-SCOPED: the box can delete its LIVE repository but cannot write to the daily snapshots of it,
so a deletion costs at most one day plus a per-file recovery - not open-ended loss. The ranking is
Viktor's; it has been #1 since July on the old story.
R-432 FILED: a sub-account sees /.zfs/snapshot EMPTY while the same box holds seven snapshots, so
per-file recovery is operator-only today. One panel read settles whether a NAMED snapshot can still be
entered, which would make it product-reachable.
R-431 SHIPPED. Third signal in OffsiteChecker. On the hub deliberately: a detector on the box is one
the deletion can silence. Threshold REASONED, not invented - over 12 898 reports every decrease lands
on ZERO and predates stats_known, and in the stats_known window there are none, so observed churn gave
nothing to calibrate against. Retention cannot halve a total; a mass deletion goes to ~0. Hence: more
than half, and at least 5. Guarded by StatsKnown (R-331), the declared State (R-204) and run success
(R-100 - whose lesson lives in this very file).
ACCEPTANCE: 9 009 real report points replayed through the detector produced ZERO alarms.
Three red-proofs run. The escalation one only became real after the first version was found HOLLOW -
it re-swept the same report, so the baseline had already moved and the latch was never consulted.
07 row 10's status is NOT moved: the write-refusal is measured, but the recovery ROUTE has never been
walked, which is what PARTIAL means.
|
||
|
|
0476a8d8e6 |
SPIKE R-95: the safety net cannot be seen from the box - and that RAISES the urgency
gates / gates (push) Successful in 17s
READ-ONLY STUDY. No code, no version, no image, no golden. No delete verb was issued against any live store. ep0, DooPlex and Peti's box were not touched at all. Q1 FIRST, AND IT DID NOT GO THE EXPECTED WAY. The brief supposed a working seven-day snapshot net might bound the worst case. Measured on BOTH boxes over their own SFTP credential, with a positive and a negative control on each: NO .snapshots is visible to either sub-account - not in the account home, not inside the repo - and the account is jailed at /. Either none exist or a sub-account cannot see them, and the second is not a reprieve: a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel. The register's claim rests on nothing that was checked. The row has NO R-number, so nothing can cite it; its "confirm tomorrow" was 2026-07-27, 36 days ago; and the DUE-CHECKS block built for exactly this (R-341) is EMPTY. R-95's word "ARMED" is withdrawn pending R-429. The confirming field is a Hetzner API field, so this spike STOPPED at the section 11-D fence and left it for Viktor - ten minutes in the panel, and it re-ranks everything. Q2: TEN verbs, not nine. `check` was missing from the brief's list; `dump` is not a verb (it is a progress phase constant) and was withdrawn. There are TWO `forget --prune` sites - offbox.go:1388 AND offbox.go:1759 - and disarming one without the other reproduces R-191 exactly. Q3 (documented, from our own API mirror): AccessSettings has five booleans and `readonly` is the only permission axis. No append-only. So the PBS shape does NOT transfer - PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing. Q5 (measured, faithful append-only model, both controls passed first): withdrawing delete does NOT wedge the store - restic treats a dead owner's lock as stale and proceeds. The constraint everyone feared is not the blocker. But `unlock --remove-all` printed "successfully removed locks" while the lock was still there, and resticStep's crash-lock self-heal is built on that call - R-430. Q6 (measured, with a control): restic 0.14.0 DOES speak rest:. Append-only is a rest-server flag, not a restic one. Q7 (measured): detection is nearly free. snapshot_count already reaches the hub and the hub APPENDS reports, so the history to compare against is already on disk. RECOMMENDATION: answer Q1 today (Viktor, ten minutes), then build detection, then move retention off the box. Defer the transport change until Q1 is answered. Register: OPEN 179 -> 181. Filed R-429, R-430; R-95 updated and kept OPEN. 07 row 10's status is deliberately UNCHANGED. |
||
|
|
2d88776227 |
R-428: the gate that hunts label-matching was matching a label (and CI proved it)
gates / gates (push) Successful in 20s
decoy_coverage_gate.py identified a repository by os.path.basename(root), looked up in a RUNNERS map. Gitea's act-runner checks the repo out into a directory called `hostexecutor`, so on its FIRST CI run the gate reported "unknown repo 'hostexecutor'" and went INCONCLUSIVE - correctly refusing to pass, and blind. A NAME standing in for a FACT, in the first ten lines of the main loop of the gate written that same morning to catch exactly that, by a session with the four shapes on screen. That is the point of R-428 and why it is recorded rather than quietly patched: this class is not carelessness. FIXED: the repo is now identified by which registered runner FILE exists under the root. Verified under a renamed directory - 14 gates found where the name-based version found none. CI also now fetches app-catalog-felhom.eu. The meta-gate walks all four runners, and a gate that cannot see part of its subject must not report a pass on it - the same reasoning, and the same fix, as the two sibling fetches already in the workflow. NOTE ON THE OTHER THREE RED RUNS, all mine and all ordering: felhom-agent and felhom-controller cite R-421, and I pushed them BEFORE felhom.eu carried that row, so instructions_gate correctly convicted "cites R-421, which appears in neither register". The register lives in felhom.eu; any repo citing a new row must be pushed after it. Re-run below. |
||
|
|
94555614ab |
observations_gate: normalise emphasis before matching a marker (R-419 follow-up)
gates / gates (push) Failing after 18s
My first fix was too strict. It anchored a marker to a line start or a bare '. ', which misses the commonest real shape - a bolded sentence followed by a bolded marker: '...them.** **FILED: R-427**'. CAUGHT BY THE GATE CONVICTING THE VERY REPORT THAT DOCUMENTS IT, on two of its own observations. That is the both-directions check working: a gate that rejects the decoy AND the genuine article is worse than the hole it replaced. Code spans are stripped FIRST and that order matters - a marker inside backticks is being talked about, never used, and removing it is what makes the R-419 decoy fail. Emphasis is stripped second so '**FILED: R-1**' and 'FILED: R-1' are the same thing to the anchor. Decoy suite re-run: the R-419 decoy and a backticked-only mention are still REFUSED; both genuine marker shapes pass. |
||
|
|
574f5df107 |
the decoy sweep: 29 gates read, 16 fooled, 10 fixed - and a gate that refuses the next one (R-421)
gates / gates (push) Failing after 17s
THE CLASS, now a row: an instrument that matches a LABEL rather than the fact it names. Five instances - R-410, R-400, R-378, R-419, R-94 - and EVERY ONE was found by accident, by someone looking at something else. The gates enforce every other rule in this project, including the rule that findings must be written down rather than left in prose. Nothing had ever checked the gates. METHOD, and it is the transferable part: for each gate, construct the label WITHOUT the fact - a directory with the right name and no bake log, a handler case that exists only in a comment, a note whose prose mentions the marker it lacks - run the gate, record what it says. No verdict was reached by reading. Reading is how all five hid. RESULT: 29 distinct scripts (35 registrations; three are shared across three runners). 19 sound, 4 holes left OPEN with rows, 6 that no plausible decoy could be built for and are named UNTESTED rather than called sound. A gate nobody tried to fool is UNKNOWN. SCOPE IS A FACT TOO - the largest single cause, and mundane. Eight gates decided what to look at with os.listdir, one level. Every one was green AND CORRECT today, and every one would have gone blind the moment anyone added a subdirectory. mojibake and docker-v already used os.walk, caught the identical planted file, and are the control that proves the cause was the listing and not the decoy. IN THIS REPO: hub-confirm and manifest-bearer now walk. observations_gate (R-419, CLOSED) requires a marker at a line start or after a sentence boundary and strips inline code spans - a note SAYING it carries no marker no longer satisfies the marker test. closed-register now CONVICTS on a row it cannot parse instead of warning: FOUR rows were in that state, TWO of them written by the session that closed them the day before, and every one was exempt from the only check that reads that file. The rows were repaired first and the conviction added second - registering a failing gate refuses every push. THE META-GATE: decoy_coverage_gate.py refuses a gate registered without a decoy or a named exemption. It convicted ITSELF the moment it was registered, which is how it came to have one. Coverage is a DECLARATION the gate AST-parses, never a grep - searching a test file for a gate's name would be the very shape this sweep exists to find. The 20 uncovered gates are listed by name (R-426). NOT FIXED, each with a row and a decoy asserting TODAY's behaviour so the fix must be deliberate: R-422 reuse-refs (only 7 extensions; a rotted .md citation is invisible), R-423 site (PAGES is a hardcoded list of 7), R-424 one-register (a defect parked as `idea`), R-425 offbox-rename (fixed FILES list). R-427: closed_register_gate checks ONE direction - twelve open rows carry a closed verdict and were NOT moved, because telling finished from partly-finished is a judgement and R-378 is the record of a machine getting it wrong. FIVE DECOYS WITHDRAWN AS ILLEGITIMATE, mine, named in the audit. A decoy nobody would write proves nothing, and manufacturing a finding to fill a row is worse than an honest NO. No product code. No version bump. No image. No golden owed. All four runners green. Register: OPEN 172 -> 178, CLOSED 160 -> 161. |
||
|
|
1e6c387a0b |
R-404 CLOSED with the ruling; R-417 CLOSED by cause removal; R-418/419/420 filed
gates / gates (push) Successful in 17s
THE RULING WAS NEITHER OPTION AS FRAMED. Both offered answers - narrow the gate, or leave it and write waivers - argued about the gate, and the gate was never the problem. DIAGNOSIS, from live source: golden_currency_gate.py never looks at the push. It compares the controller's newest CHANGELOG heading against this repo's bake evidence and returns the same verdict whatever you are pushing - correct for a standing invariant, wrong as a push gate. And controller_gates.py had NO golden-currency entry at all. So the repo where a release happens never checked, and the repo that cannot create the debt was refused on every push. 18 of the last 24 pushes here touched no code - measured, and the new classifier agrees EXACTLY - most of them by construction, because the controller's code is in one repo and its register lives in this one. SIX of those 18 were bake records, so the push that PAYS the debt is itself documents-only: the gate was blocking its own cure. Not the waiver its docstring prescribes: that clause was written for a release nobody wants a golden for. R-417 was a release we DID want a golden for, on a night the runbook forbade baking. A waiver would have recorded a lie. RULING: block the push that can create the debt, notify the push that cannot. The gate's logic, exit codes and wording are BYTE-IDENTICAL. Only the consequence changed, for one gate, on one kind of push, with a loud ADVISORY block so nothing goes quiet. R-242's vouch half is amended in place to say it is UNTOUCHED and still open - a baked-but-unvouched golden still passes both the gate and the new notice. Do not read R-404's closure as closing it. FILED: R-418 - this runner's docstring listed ELEVEN gates while THIRTEEN were registered; one-register and closed-register ran undocumented since 2026-08-24. Enumeration fixed here, the correspondence is still unenforced. R-419 - observations_gate.py accepts an item whose body merely CONTAINS "NOT-A-FINDING", even in prose disclaiming it; found by accident when a planted test observation passed and my live validation proved nothing. R-420 - controller_gates.py could not express a non-blocking gate at all before today. Register: OPEN 171 -> 172, CLOSED 158 -> 160. |
||
|
|
1f74427fd2 |
target-selection: a drill night will see the golden ADVISORY, and that is expected (R-417)
The instruction that produced R-417 was a prompt, not a file, so the next drill author would have met the same surprise. This is the durable home they actually read before picking a machine. Says what to expect (a loud ADVISORY on every documents push, all night), what still refuses (code pushes, and every other gate), and what the honest instrument is if a release must ship without a golden - a waiver row, never --no-verify. |
||
|
|
1c00af607c |
R-404: block the push that can create the golden debt, notify the one that cannot
push_scope.py classifies a push as code or documents from an ALLOW-LIST of document paths - everything else, including any new top-level directory, is code. Every uncertainty (first push, force-push, merge commit, empty range, unreadable stdin) answers code: guessing 'documents' would hand out the exemption by accident. repo_gates.py gains a fifth GATES field and --scope=code|docs. On a documents-only push a golden-currency CONVICTION prints as ADVISORY in its own block and does not refuse; every other gate still refuses every push, and golden-currency still refuses a push touching code. The gate itself is UNCHANGED - its verdict, exit codes and wording are byte-identical. What changed is who is refused. Measured on git 2.47.3: a pre-push hook receives <local ref> <local sha> <remote ref> <remote sha> on stdin, one line per ref; a first push carries an all-zero remote sha and a deletion an all-zero local sha. Both land on code. |
||
|
|
a91c0580eb |
R-417: the golden gate and a drill night cannot both be satisfied; two instruction defects fixed
gates / gates (push) Successful in 16s
Five felhom.eu CI runs went red tonight (jobs 469/470/471/473/476) and all five were mine, every
one on step 3 `Run the gate entry point`. Job 478 is green. CAUSE CONFIRMED BY ISOLATION: the only
functional diff between the last red and the green is the golden-0.232.0 evidence directory;
moving it aside reproduces exit=1, restoring it gives exit=0, tree byte-identical after. The first
reproduction attempt used a detached worktree, where three gates go INCONCLUSIVE for want of the
sibling clones - that is the worktree, not the commit, so it is discarded rather than quoted.
THE GATE WAS RIGHT EVERY TIME. 0.231.0 and 0.232.0 were released with no golden carrying them, so
a machine installed in those hours would have received 0.230.0.
R-417 is the SHAPE, not the gate: the soak runbook forbade baking a golden that night, so red was
unavoidable and pushing the drill's own evidence needed --no-verify. The gate's failure text names
the remedy for exactly that case - record a waiver here, never a bypass - and I did not write one.
A red CI run on a drill night is now indistinguishable from a real one, which is the whole value
of the signal.
TWO INSTRUCTION DEFECTS, found by following the end-of-session checklist and being unable to:
- The CI-verification recipe cannot produce what it asks for. `actions/tasks` returns
"conclusion": null for every run, so a session following it quotes a conclusion it never read.
Its id is also offset from the `jobs` id for the same run (479 vs 478 for
|
||
|
|
63eff21a5c |
golden 0.232.0 baked, vouched, floor raised — and it carries TWO releases
gates / gates (push) Successful in 17s
GOLDEN_SHA256 5f8a53ed5b19a6cb2006298ce6239f6fca2b990cc3ef6eada89f602801ca91b8, 657 494 489 B. 0.231.0 was never baked, so the fleet went 0.230.0 -> 0.232.0. THE CHECK THE 0.230.0 BAKE SKIPPED, AND THIS ONE DID NOT: the bake script's fingerprint was compared ACROSS THE HOP - 7b0fb5cf...73b6a1 on DooPlex and inside the VM. The previous bake recorded only the DooPlex-side hash and said so; this one is a measurement. Three independent readers agreed before anything was vouched: the bake's own print, the round trip of the published bytes (HTTP 200, 657494489 B, same sha, hashed from what was downloaded), and the hub's Day-0 dropdown reading Gitea on a different code path. The delivered artifact names its own controller - ./etc/felhom-controller-image reads felhom-controller:0.232.0 - with 19382 entries under var/lib/felhom/docker/. Both pre-gates were shown able to see something before their zeroes were believed, and the manifest was RE-READ after vouching rather than trusted from the 303 flash. DELIVERY WAS ACTUALLY EXERCISED. Both boxes had been hand-deployed during validation, so the floor had nothing to move. Rather than report delivery untested, demo-felhom was rolled back to 0.231.0 and the chain run for real - it moved itself in ~20s: 10:47:16 controller-swap: image file written, restarting bootstrap target=...0.232.0 10:47:26 controller-swap: new controller healthy target=...0.232.0 And this is the first bake golden_currency_gate.py actually gates: it now reads the GOLDEN_SHA256 line out of the bake log rather than matching a directory name (R-410, shipped hours earlier the same day). All 13 felhom.eu gates are green, golden-currency included, for the first time since v0.230.0 was released. |
||
|
|
f41a1a0ad8 |
R-411/408/407, R-414, R-412a leg 1, R-410, R-406 CLOSED; determination + live evidence
gates / gates (push) Failing after 18s
Part 2.1's determination is the first artifact: the scratch resolver was consciously OUT OF SCOPE for R-356, not excluded on state-only grounds - established from R-356's own commit 08eb1a6, whose tests say 'the prepared scratch still resolves ... only the DESTINATION moves'. So 07 section 6.3's rule applies and now has a FOURTH consumer, and the section says so. Live evidence: the collision rerun on demo-hp with the sampler positively controlled first (12 locks=1 across a real check, 4 locks=0 quiet), showing unlock --remove-all 0 times where the drill saw it twice; and the proof reaching verdict pass on demo-felhom - the box that could not run it at all - recorded where last_proof_result had been ABSENT every night. Capability map: the off-site proof row now records that the nightly firing IS proven (it ran unattended at 05:30 on demo-hp) and that a driveless box can now be proved. Register: six rows closed and compressed. OPEN 176 -> 170, CLOSED 152 -> 158. |
||
|
|
22e1c95e6a |
the golden gate reads a fact not a name (R-410); R-133's collision resolved (R-406, R-416)
gates / gates (push) Failing after 17s
R-410. golden_currency_gate.py matched EVIDENCE_RE against os.listdir and read nothing inside, so `mkdir documentation/tests/golden-9.9.9-2026-01-01` turned it green with no bake behind it - noticed while the 0.230.0 bake was running, when the evidence directory existed before the bake finished. It now reads the GOLDEN_SHA256= line out of that directory's bake log: a directory name is a label, that line is a fact only a completed publish produces. Still offline, still --fast, one file read. Directories that look right and hold nothing are printed by name rather than silently ignored, so a half-finished bake is visible. test_golden_currency_gate.py ships the red-proof with a POSITIVE CONTROL, without which "it fails on an empty directory" would be satisfied by a gate that fails on everything: CASE 1 empty directory -> rejected and named CASE 2 log with no GOLDEN_SHA256 -> rejected CASE 3 real bake log -> counted, and its sha read <- the control CASE 4 the tree is left byte-identical Red-proofed: reverting the gate to name-matching fails cases 1, 2 and 3. R-406. Citations MEASURED before choosing, which is what the row asked for: hub-uniqueness had 3 references (all inside one audit doc), plaintext-break-glass had 5 (CONTEXT.md, break-glass.md, hub/CHANGELOG.md, the capability map, a spike). The FEWER-cited one moved - hub uniqueness is now R-415 - and all three citations were rewritten to "R-415 (was R-133)" rather than silently swapped. THIS IS THE OPPOSITE OF THE TASK'S LITERAL INSTRUCTION, which said renumber the second row on the stated ground that "the older number has the longer reference trail". Measured, that ground points the other way. The principle was followed and the letter was not, and the row says so rather than leaving an unexplained diff. R-416 filed: the within-register duplicate rule was deliberately NOT added in the same commit that removed its only subject - a guard whose red-proof can only be a planted fixture is not this project's standard. Now that the register is clean it can ship with the next real duplicate as its first subject. |
||
|
|
f8f9ffdf2b |
STATUS: R-414 is the morning's first item - the new nightly check cannot run on demo-felhom
gates / gates (push) Failing after 17s
|
||
|
|
cee8f70e98 |
soak drill COMPLETE: 7 phases, 4 findings, the observer found the worst one (R-414)
gates / gates (push) Failing after 17s
Overnight soak 22:39->06:10 CEST. demo-hp the victim, demo-felhom the untouched observer.
No production code, no golden, no version bump. Report at
documentation/audits/DRILL-soak-2026-08-31/REPORT.md.
VERDICTS: 1 lock-collision FAIL, 2 guard-interactions PASS-with-one-defect, 3 R-357 PASS,
4 proof-edges PASS, 5 mutated-cycle PASS, 6 observer FAIL, 7 teardown PASS.
R-414 - THE MOST VALUABLE FINDING, AND ONLY AN UNTOUCHED BOX COULD HAVE FOUND IT. On
demo-felhom the nightly proof fired for the first time unattended at 05:30 and REFUSED:
"nowhere to restore to - nincs regisztralt adatmeghajto". Cause established, not inferred:
storage_paths is EMPTY, so there is no path to put a scratch on. It will fail this way every
night forever with only a WARN, and because the error path reaches no verdict,
last_proof_result stays ABSENT - which is also what a pre-0.231.0 controller sends. The hub
cannot tell "never ran" from "not deployed": the StatsKnown trap one level up. The box is
NOT unprotected; its off-site backup ran fine in 46.9s. It is the PROOF that cannot run.
R-411 - measured, not reasoned: restic stats TAKES A LOCK; a customer full-restore runs it
while holding no acquireRunning; the integrity check is therefore not blocked, meets that
lock and escalates to unlock --remove-all. The sampler caught "restore ..." and
"unlock --remove-all" in the SAME sample. Contained: the check was classified unreachable,
not damage, so no false alarm.
R-412 - CORRECTED from HIGH to LOW. I filed it on a mechanism I had not finished measuring.
The off-site run has its own pre-push dump leg, so a hollow unit is REPAIRED before it
ships - proven on two apps and confirmed by pulling the snapshot back out of the store.
What survives is a narrow race, plus a success line over a backup holding no data.
R-413 - the R-87 proof caught a product-produced hollow snapshot unattended, and the
nightly job fired on its own schedule at 05:30 for the first time (bentopdf PASSED on
9d002b38 in 2.315s). Both were listed "not yet live-validated" yesterday.
R-403 mirror guard PROVEN live, with a negative control: it fired when a unit was hollow
("The copy was PRESERVED rather than replaced with an empty one") and skipped 0 legs at
teardown when every unit was sound.
R-357 PASS at last, six days owed: a real full filesystem, refused BEFORE StopStack, app
never stopped, live data byte-identical, and it worked once the space came back.
Phase 4 built the false-alarm control the whole R-87 design rests on: bentopdf is the only
template of 53 with neither a database nor a named volume. It passes silently.
EIGHT of my own instrument errors are named in the report, each caught by its own control -
including a time guard that fired an injection four hours early, and filing R-412 at the
wrong severity.
Teardown clean on all three layers of both boxes; both healthy on 0.231.0.
OWED: a golden for 0.231.0, and a decision on keeping bentopdf.
|
||
|
|
ab8b884763 |
soak phase 5: R-403 guard PROVEN live; R-412 CORRECTED down after measuring the mechanism
gates / gates (push) Failing after 18s
R-403 MIRROR GUARD - PASS, forced after the natural test evaporated. privatebin was injected hollow at 23:34 to meet the 03:30 mirror; the 02:30 db-dump re-made its tar, so by 03:30 the primary was complete and the guard had nothing to refuse. Forced instead on calibre-web through the real Tier-2 path: the guard fired and named itself - "unit leg SKIPPED ... The copy was PRESERVED rather than replaced with an empty one (R-403). The other legs continue." Secondary byte-identical, 23 files, tar sha d7e7f422. R-412 CORRECTED, AND I OVERSTATED IT WHEN I FILED IT. The first wording claimed the hollow unit sits in the store for a whole cycle because the volume-dump leg runs only on the backup schedule. That is WRONG. The 04:15 off-site run has its OWN pre-push dump leg - "Stopping calibre-web for safe volume dump", "Volume dump: ... -> 877.5 KB" - so a unit that is hollow when a run starts is REPAIRED before it is pushed. Measured twice: opengist and calibre-web both went in hollow and came out complete, and the snapshot pulled back from the store (6fee3b5a) holds the volume tar and all 17 userdata files. What remains real is narrower: the one hollow snapshot that DID reach the store was created when the unit was destroyed INSIDE a run that had already completed that app's dump leg. The race is real and was observed, and "backed up opengist (... 0 mandatory path(s))" is a success line over a backup holding none of the app's data either way. Severity HIGH -> LOW, with the correction stated in the row rather than quietly rewritten. Phase 6 interim: the observer is clean so far - db-dump 674ms, tier2-backup 3ms (a no-op, cause to be established not assumed), zero ERROR/WARN since 23:00. Also recorded: two of Phase 5's four injections were NOT performed, with the reasons established rather than asserted - there is no endpoint that reaches SetDisconnected and a hand-set flag would be reverted by the live monitor before 04:15; and a corrupted manifest provably never reaches the store because the capture rewrites it first. |
||
|
|
585ed654b4 |
soak drill phases 0-4: R-408's hazard is REAL and reachable; R-357 finally live (R-411..R-413)
gates / gates (push) Failing after 17s
Overnight soak on demo-hp, phases 0-4 of 7. Evidence documentation/audits/DRILL-soak-2026-08-31/. No production code written, no golden, no version bump - findings only. PHASE 1 FAIL - R-411. The R-408 hazard was only ever reasoned about; tonight it was measured through the product's own endpoints. restic stats TAKES A LOCK (clean-room: nothing else running, 4x stats, sampler reads locks=1). A customer full-restore runs stats in its preparation while holding NO acquireRunning, so the integrity check is not blocked, runs, meets that lock, and resticStep escalates to unlock --remove-all - the sampler caught "restore ..." and "unlock --remove-all" in the SAME sample at 20:50:51. The log calls it "a stale exclusive lock left by a previous crash"; there was no crash. THE CUSTOMER-FACING CONSEQUENCE IS CONTAINED and that is R-359 working: the check was classified Unreachable, NOT damage, so no alarm and due-ness held. The opposite direction is FENCED - five restores fired into a running check at 5/15/25/35/40s were all refused by restoreOpBlocked with zero restic invoked. PHASE 2 FAIL - R-412, and it is the natural instance of R-403's shape that yesterday's session had to hand-build. A lost primary unit is rebuilt by the capture WITHOUT its volume dumps (185664 B -> 4382 B, volume_dumps: None), and the off-site backup then ships it and logs "backed up opengist ... 0 mandatory path(s)" - a success line over a backup holding none of the app's data. PHASE 2 also R-413: the R-87 proof CAUGHT that hollow snapshot unattended - verdict "fail", volumes_expected_none_captured: opengist_data, one offsite_proof_empty at severity error, while the four apps ahead of it in the rotation passed. 2.3 stale marker PASS, 2.4 two restores PASS, 2.5 corrected the runbook's premise (RunTier2 has zero acquireRunning and zero restic references - a local mirror cannot collide with the repo, so the proof correctly does not skip for it). PHASE 3 PASS - R-357 live-validated at last, six days owed. A real full filesystem (1900544 B free vs 5878421 B needed) on a separate 1 TB device that does not back Docker. Refused BEFORE StopStack: app "Up 6 minutes (healthy)" unchanged, 0 safety dumps, live userdata tree fingerprint 5a9b2db75dfd4ba4131adb3255670472 unchanged, message names both numbers in Hungarian. Ballast removed, the SAME restore then worked and the fingerprint is still identical. On the same full disk the Tier-2 run, the proof and the integrity check all behaved: the proof refused before any download through the shared unitOnlyHeadroom gate extracted today, reached no verdict and did not alarm. PHASE 4 PASS - the false-alarm control the whole R-87 design rests on now EXISTS. Found by reading all 53 catalogue composes: bentopdf is the only template with neither a database service nor a named volume. Deployed, backed up, proved: PASSED in 2.224s with zero alarms. Four of my own instrument errors were caught by their own controls before any result was believed: a hub log line used as a controller positive control, a grep pattern that missed a registered job, a heredoc that mangled a planted marker, and a catalogue scan that read 0 of 0 apps from the wrong path. Phases 5-7 to follow: the real nightly cycle with injected faults, the untouched observer, and teardown. |
||
|
|
7ee25925f9 |
R-87 CLOSED: live evidence, capability row, architecture verdict, registers
gates / gates (push) Failing after 17s
Controller v0.231.0 + hub v0.110.0, both deployed and verified on demo-hp. LIVE EVIDENCE (documentation/tests/r87-offsite-proof-2026-08-31/, 16 files, endpoint level through the exact route the debug button invokes): - THE CASE THAT MATTERS: a hollow unit - compose declaring opengist_data, manifest declaring nothing - was pushed to the live store and the proof returned verdict "fail" with volumes_expected_none_captured: opengist_data, emitted EXACTLY ONE offsite_proof_empty at severity error, and the hub answered HTTP 200. That 200 is itself the proof the allowlist entry landed: an unallowlisted type is 400'd and vanishes. - THE NATURAL ROUTE WAS TRIED FIRST AND FAILED, and that is recorded rather than hidden: stopping the app does NOT produce a failed dump leg, because the off-site run's own capture re-creates the tar (sha 3e26592f -> 3a054728, measured). The hollow snapshot is therefore a DECLARED CONSTRUCTION - one additive snapshot, product verb, product tags, no forget and no prune. State restored: the product's own run made a healthy snapshot the newest again and the proof then passed opengist. - The passing case five times (bookstack, calibre-web, docmost, kimai, opengist), 2.2-4.0s each, matching the spike's measured band. - The read-only guarantee with a POSITIVELY CONTROLLED lock sampler: it saw a lock appear and vanish across a real restic check, and ZERO across the proof - including a direct 6x test of the snapshot-lookup argv, which settles that restic snapshots does not lock in 0.14.0 either. - Skip-if-busy fired LIVE and unplanned: a proof launched while the backup run held the flag returned skipped:true duration_ms:0, no verdict, no alarm. - The customer's own verification copies were untouched throughout, which is the safety property the separate proof root exists for. ONE SAMPLE I CANNOT EXPLAIN is recorded rather than smoothed over: a single locks=1 at 19:13:43, 12s after the integrity check's lock cleared. Two independent tests exclude the proof; I did not establish what it was. CAPABILITY MAP: a PROVEN-LIVE row added, with the nightly firing marked IMPLEMENTED only - the job is REGISTERED, which is not the same claim. 07 section 8 MATRIX ROW 4 WAS NOT MOVED, deliberately, and section 10.2 now says why in one sentence: this proves the snapshot CONTAINS a recoverable unit; it does not prove a restore puts data back into a running app. Without that sentence the new green tick reads as covering the drill. REGISTER: R-87 CLOSED and compressed into CLOSED-ITEMS.md. OPEN 172 -> 171, CLOSED 151 -> 152. No new rows minted. R-408 and R-409 stay open and are referenced by this work. golden-currency is RED and it is a DECLARED, EXPECTED debt: v0.231.0 is released and the newest golden carries 0.230.0. The fleet is on 0.230.0; demo-felhom does not have this job. A golden carrying 0.231.0 is OWED and it is Viktor's call (R-242). This push uses --no-verify for that reason - bypass #8. |
||
|
|
1aeaa30c28 |
hub v0.110.0: allowlist + operator-only for offsite_proof_empty (R-87)
gates / gates (push) Successful in 15s
Controller v0.231.0 adds a nightly job that proves an app's newest off-site snapshot still CONTAINS that app's data. When it finds one that does not, it emits offsite_proof_empty. Two register lines, both load-bearing and both in this commit: - allowedEventTypes - an unallowlisted type is answered 400 and VANISHES, so this entry is what makes the alarm exist at all. - operatorOnlyEvents - a missing customerMessages entry is NOT a routing block (FormatCustomerEmail falls back to the raw message, the v0.78.0 defect that register was built for). A customer can take no action on a hollow recovery unit. DELIBERATELY NOT reusing backup_integrity_failed, which is the nearest existing type: it means THE STORE IS DAMAGED and carries the Hungarian template saying so. Here the store is sound and the CONTENT is absent - different cause, different action, and telling a customer their backups are damaged when they are not is the more expensive mistake. Same asymmetry looksLikeRepositoryDamage is shaped around. DELIBERATELY no customerMessages entry (the controller's dynamic Hungarian names the app and what is missing; a template would discard it) and DELIBERATELY not in perAppCooldownEvents (a fenced act - this job proves ONE app per night, so the coarse hourly cooldown is already the right grain). This widens the R-87 task's stated repo scope to felhom.eu/hub/. The reason is recorded in felhom-controller/CONTEXT.md ruling 4 rather than left as an unexplained diff. Hub green gate: go build/vet/test all pass, 18 packages. |
||
|
|
177c75781e |
REPORT: CI run 285 is GREEN on the golden-bake commit - the golden debt closed the gate
gates / gates (push) Successful in 17s
|
||
|
|
2263245cf2 |
golden 0.230.0 baked, vouched, floor raised - demo-felhom moved itself off the R-403 build (R-410 filed)
gates / gates (push) Successful in 17s
GOLDEN_SHA256 9287f7cef5f13166276e8406005e3f28004004510c5184f1c1c7377f7aafad2e, 657 873 700 B. Evidence documentation/tests/golden-0.230.0-2026-08-31/. WHY IT WAS OWED: the newest golden was 0.229.0, which IS the build R-403 says deletes a good copy. Every fresh install and the whole fleet floor still carried it. golden_currency_gate.py had been red across |
||
|
|
32a4c35c9c |
REPORT: record the CI verdict by run id - 283 red on the pre-existing golden debt, and 281 proves it is pre-existing
gates / gates (push) Failing after 17s
|
||
|
|
130f7a6eba |
R-87 SPIKE: measured, do not build it as written (R-407..R-409 filed)
gates / gates (push) Failing after 17s
Spike. NO production code. No version bump, no build, no deploy, no golden.
felhom-controller and felhom-agent were READ ONLY. The fleet stays on v0.230.0.
Q1 restic is 0.14.0 (go1.19.8, bookworm 12.15) - the four source comments asserting
it are CONFIRMED, not corrected.
Q2 --verify DOES exist and is NOT a content check. Red-proof: one byte changed in a
restored 160 MB tar with size and mtime preserved passed clean, rc=0. Verify took
131 ms on a 213 MB / 7-file tree, which cannot be hashing. A size or mtime mismatch
causes a silent re-download, not a failure. Controls: --target 1 hit, four post-0.14
flags and a nonsense string 0 hits each. Neither --verify nor --no-lock appears
anywhere in the controller source.
Q3 no reference for "correct" exists. restic ls --json carries no content hash in
0.14.0, and the unit manifest hashes 4918 B of a 213231242 B unit - 0.0023 percent,
the config files and not the dumps or the tars. R-409.
Q4 it is CHEAP. All 8 apps / 774378123 B logical restored back to back in 25 s, against
40257 ms for the weekly 100 percent check beside it. Individual restores 2253-3978 ms
regardless of size: cost is per-snapshot round-trip plus ~1 s per 200 MB. Peak scratch
is the app's full logical size. The 1.1 MB restic cache is index only and hides nothing
(--no-cache 5423 ms vs cached 3198 ms, trees byte-identical).
Q5 skip-if-busy stays right at 25 s against a 2m52s nightly backup. But
RestoreOffboxScratch takes NO acquireRunning, while offbox_integrity.go:28 asserts
every off-site operation does. R-408.
Q6 observed with a positively-controlled lock sampler: restic restore takes NO lock;
restic check DOES (locks 0 -> 1 for nine samples -> 0 across the check, zero across two
restores). The product writes anyway - unlockStale runs `restic unlock`, a delete verb,
before every restore (offbox_restore.go:289). The task's lead was right in direction and
wrong in mechanism. R-95's constraint IS satisfiable: --no-lock plus skipping unlockStale
writes nothing, and both mechanisms exist unused. Neither was fixed - the task forbids it.
offbox_integrity.go:255's "It NEVER writes to the repository" is R-407.
Q7 THE DECIDING ONE: of R-353/354/356/358/403 an unattended scratch-restore would have
caught ONE (R-356). The value is elsewhere, and the weekly check structurally cannot
reach it: `check` proves the stored bytes are the stored bytes, never that we stored the
RIGHT thing. A hollow unit backs up, checks at 100 percent and restores cleanly and
recovers nothing - R-403, measured in bytes on 31 August.
RECOMMENDATION: option C, the NARROW test - one app a night, restored to scratch, checked
against its own manifest.json through the existing unitCarriesData, scratch deleted, the
SNAPSHOT recorded as the proof. Options A (do not build) and B (scheduled attended drill)
considered explicitly; B is weakest because it is what already happens. R-87 should be
RE-SCOPED, not built as written, and that is Viktor's call - the row stays open carrying
the verdict and STATUS.md item 4 asks it in plain words.
Also corrected in 07-backup-architecture.md: matrix rows 4 and 10 both said "the depth
that ships ON does not re-read pack contents (R-399)". R-399 CLOSED in v0.228.0 and the
depth is 100 percent. Two stale cells, fixed, and the spike verdict added beside them.
Row 4's verdict is UNCHANGED by the spike and now says so.
Teardown: all three layers, none of them "nothing was created" - 6 files on the PVE host,
9 in the guest, 5 plus 2 run-flags in the container, all removed and verified empty. The
four scratch directories this session's restores created were removed; three that
pre-date the session were left alone. Two state changes recorded rather than hidden: the
control integrity run recorded its verdict (depth structure -> 100%, due-ness +7 days),
and four restores appear in the controller log. Nothing was written to the off-site
repository by hand.
Evidence: documentation/audits/evidence-spike-restic-restore-2026-08-31/ - 31 files,
every one pulled off the box BEFORE teardown (R-320).
golden-currency is RED at this commit and was already red at
|
||
|
|
6e550aedd3 |
R-87 put back in the register; closed_register_gate.py is the 12th gate (R-405, R-406)
gates / gates (push) Failing after 17s
Records and process only. No machine contacted. No product code, no version bump, no build, no deploy. R-87 was moved into CLOSED-ITEMS.md by the 2026-08-22 compression sweep |
||
|
|
dddcc808be |
R-403 CLOSED (controller v0.230.0), R-404 filed as a decision for Viktor
gates / gates (push) Failing after 17s
07-backup-architecture gains section 8.2, placed beside row 5 on purpose: the derived-copy rebuild rule is UNCHANGED and section 8.2 names the single exception, so a future reader who finds RunTier2 skipping a leg does not fix it back. It carries the measurement (120 082 104 B -> 7 036 B on the shipped v0.229.0), the four-case table, why hollow is a manifest question and not a size question, why the data legs are deliberately not guarded, and why the capture job is not guarded either. 00-capability-map: the Tier-2 row's status does NOT move, stated explicitly rather than left ambiguous. R-403 removes a way the route could be DESTROYED between uses; it does not change what the route can be relied on for. Register: R-403 CLOSED and compressed into CLOSED-ITEMS (594 -> 593 open lines). R-404 FILED as a DECISION and deliberately not acted on - should a documents-only push be subject to the golden-currency gate, now that it has been correctly bypassed six times? Both sides stated, plus what happens if Viktor does nothing. The gate was NOT changed. R-242: seventh conviction, and the FIRST where the day-0 ground does not apply - R-403 is a defect in the nightly Tier-2 copy, which a day-0 box starts running on its first night. This push uses git push --no-verify, declared here and in felhom-controller/REPORT.md. A golden carrying 0.230.0 is owed and is more urgent than the previous six. STATUS: the R-403 item moves out of 'Broken' into what works, in plain words; the delivery item now says a golden is owed and that the fleet carries the defect; R-404 goes into the decide section with its do-nothing outcome. Drill evidence: nine phase logs, including the two things that went wrong (a repair whose rsync was not installed in the guest and silently did nothing, and a session that expired mid-run so a POST did nothing). |
||
|
|
66156c619f |
R-403 drill evidence + the credential reader that ends a three-time mistake
gates / gates (push) Successful in 16s
The drill: the loss reproduced on the shipped v0.229.0 before anything was built. 120 082 104 B -> 7 036 B in one Tier-2 run, recorded as a success. Phases 1a (before), 1b (the hollow primary, produced through the R-102 restore path exactly as the 2026-08-31 observation was), 1c (the loss), 1d (repair). scripts/read_credential.py is Part 4's rider, and it exists because a note did not work three times: 2026-07-20 a Failed login was diagnosed as a stale password and written into memory; 2026-08-31 the same misreading recurred and was caught; 2026-08-31, hours later, it recurred AGAIN and rewrote a live box's password hash. Between them the project already had a memory file stating the rule, a worked recipe in it, and a session report describing the mistake. The rule now lives in the code path: one matching quote pair is unwrapped, the result is REFUSED if it still carries a quote, and --expect-length gives the caller a second opinion. The value goes file->file at 0600 and stdout gets only its length. test_read_credential.py asserts each refusal by its reason, with a positive control before believing the not-in-stdout result. Red-proof E1: remove the final quote assertion -> three cases fail by name. |
||
|
|
83ff9e8e38 |
golden 0.229.0 baked, vouched, floor raised — R-242's sixth debt PAID the same day
gates / gates (push) Successful in 16s
GOLDEN_SHA256 39aa886df77b21757aef3b298a389343dc0df5134bb0f14e8f92a451d7bdae87, 656 864 331 B.
The evidence is the ROUND TRIP, not the build log: the published bytes were downloaded back and
match the bake on both size and sha, and ./etc/felhom-controller-image read OUT of the downloaded
archive says felhom-controller:0.229.0 - the delivered artifact naming the controller it will start.
A THIRD independent reader agreed before anything was vouched: the hub's own Day-0 dropdown read the
same sha straight from Gitea, a different code path.
Both pre-gates were proven able to see something before their negative results were believed - the
404 pre-gate against a 200 from 0.228.0, and the token-leak grep against a seeded throwaway copy.
Acceptance markers counted on the COMMITTED log: 1/1/1/1 present, 0/0/0 absent.
The vouch is a three-field change, checked rather than assumed: MinAgent 0.129.0 read from the
golden's controller CHANGELOG header, agent_version 0.130.0 >= min_agent 0.129.0 (not the R-216
shape), agent_sha256 and wrapper_sha256 carried through explicitly because the handler clears a
field it is not sent. Verified by re-reading the manifest, never by trusting the flash. The R-120
gate PASSED rather than being bypassed - fleet newest 0.229.0, golden 0.229.0.
The floor is proven ACTING, not merely set: demo-felhom self-updated 0.228.0 -> 0.229.0 and logged
settle-gate GO at/above floor 0.229.0. Nobody deployed to that box. Both demo machines now carry the
Tier-2 unit restore.
golden_currency_gate.py went red -> green; the --no-verify bypass declared on
|
||
|
|
c2de785bf2 |
R-102 + R-103 CLOSED (controller v0.229.0) — architecture, register, STATUS, drill evidence
gates / gates (push) Failing after 17s
07-backup-architecture: 6.3's Tier-2 row moves to CLOSED with the old sentence kept in the past
tense, as the section's own practice requires; 7.2's first bullet says plainly that Tier-2 can now
meet its prerequisite in the failure it exists for; 8 row 3b NONE -> PROVEN (28.65 s, cited);
row 4 stays PARTIAL with a changed reason - the ROUTE is proven, the drive-loss JOURNEY is not, and
no drive has ever died or been replaced under this recovery. 8.1's blanks updated per row.
6.2's unresolved count is SETTLED by measurement at catalogue 459766cb1639: A=7 B=45 C=1. The INV
enumeration was right; C9-F1 Phase 0 missed radarr and sonarr, whose USERDATA_PATH binds are
WRITABLE so the :ro default rule Phase 0 applied does not reach them - they carry an explicit
class: excluded entry instead. Class C is bentopdf. No catalogue file was changed.
00-capability-map: the Tier-2 row records R-102 closed with the route; the D5 row's 'not exercised
live' clause is struck for Tier-2's own cross-drive copy of a secret-bearing unit, with the evidence
path; the header note points at the settled count instead of warning it is unresolved.
Register: R-102, R-103 and their C9-F4 / C9-F1b aliases closed and compressed into CLOSED-ITEMS
(596 -> 593 lines, each naming git show
|
||
|
|
1623a4d5b5 |
golden 0.228.0 — baked, published, round-trip verified, vouched, floor raised
gates / gates (push) Successful in 19s
Bake: build-golden.sh v3.0.0 in the drill VM, reverted to virgin and cold-booted.
GOLDEN_SHA256 76a3a98b9e7cc23bf8ae51b38a6272f576df285cb34cd22235ac3f06a31e53ec,
658 079 744 B. All four acceptance markers counted 1; excluding/FATAL/mp1 counted 0.
The 404 pre-gate was proven to work before its 404 was believed — the target URL
404'd while the existing 0.227.1 package 200'd on the same command.
The evidence is the ROUND TRIP: the downloaded bytes match the bake's size and
sha, and ./etc/felhom-controller-image read OUT of the downloaded archive says
gitea.dooplex.hu/admin/felhom-controller:0.228.0.
Vouch: three fields together — golden_version 0.228.0, agent_version 0.130.0,
min_agent 0.129.0 (read from the controller CHANGELOG header, not assumed);
wrapper_sha256 carried through explicitly. agent >= min_agent, so not the R-216
shape. Verified by RE-READING the manifest, never the flash. R-120 gate passed.
Floor raised 0.227.1 -> 0.228.0 (impact preview {"below":3,"valid":true}).
The floor is ACTING: demo-felhom self-updated 0.227.1 -> 0.228.0 in under a
minute and re-registered offsite-integrity by itself. Both demo boxes now
re-read their whole off-site store on the weekly check.
Token hygiene: file->file scp, runner script inside the VM, unit properties
grepped 0. The leak grep on the committed log was proven with a planted copy
(1) before its 0 was believed. Teardown: guest 9100 purged, secrets shredded
after the log was copied out, VM off, disk reverted to virgin.
golden_currency_gate.py red -> green. All 12 felhom.eu gates OK.
|
||
|
|
77a5a1154b |
docs: controller v0.228.0 — R-399/R-400 closed, R-401/R-402 filed
gates / gates (push) Failing after 19s
STATUS.md: header said 2026-08-23 over a 2026-08-30 body, and two "Waiting on you" items were both numbered 4 — both fixed. R-399 leaves that section (decided and shipped); the depth change is stated in plain words and the remaining items each say what happens if Viktor does nothing. 00-capability-map.md: the off-site verification row now carries its DEPTH, and its live citation is the 2026-08-31 run at 100%. The weekly firing at the new depth stays IMPLEMENTED, not PROVEN-LIVE. 07-backup-architecture.md §10.2: R-399 recorded closed, with the one sentence that stops it being turned back down — the structure check PASSED a size-preserving pack corruption. R-87 untouched and still OPEN. Register: R-399 and R-400 compressed into CLOSED-ITEMS.md with their reasoning kept and 300d7e8 named as the commit holding the originals. R-401 filed with a TRIGGER (the slow-check WARN firing) rather than a date. R-402 filed: the integrity verdict and its depth are on the wire and no hub surface reads either. OPEN 166 -> 165, CLOSED 148 -> 150. wire_contract_gate.py: offsite.last_integrity_depth allowlisted WITH ITS REASON beside its sibling last_integrity_ok, both to be deleted together when a hub surface is built (R-402). |
||
|
|
db0812b6f2 |
Golden 0.227.1 baked, vouched, floor raised — and the floor delivered the new job by itself
gates / gates (push) Successful in 16s
Second full delivery of the day. golden_currency_gate.py went red -> green on the same command, so the --no-verify bypass declared on the previous push is now historical rather than standing. GOLDEN_VERSION 0.227.1 GOLDEN_SHA256 66754491dc9bd0130ef8ded9562f63c53a5ffdcfd91baa551141e55fa083ea32 size 657 403 203 B baked gitea.dooplex.hu/admin/felhom-controller:0.227.1 MinAgent 0.129.0 (read from the controller CHANGELOG header, not assumed) THE EVIDENCE IS THE ROUND TRIP. The published bytes were downloaded back -- size and sha256 identical to what the bake reported -- and ./etc/felhom-controller- image was read OUT of the downloaded archive: felhom-controller:0.227.1. That is the delivered artifact naming the controller it will start, from the bytes a customer's box would actually fetch. Markers counted: docker OK (overlay2 = 1, mount point rootfs = 1, mp0 = 1, upload OK (HTTP 201) = 1; excluding = 0, FATAL = 0, mp1 = 0. 404 pre-gate passed before the run and the script's own pre-delete agreed, so nothing was overwritten. Three-field vouch, all three checked: agent_version 0.130.0 >= min_agent 0.129.0 (NOT the R-216 shape), wrapper_sha256 carried through explicitly because the handler clears it when omitted. Verified by RE-READING the manifest rather than trusting the flash. The R-120 gate on that POST passed on its own terms rather than being worked around. AND THE LINE WORTH KEEPING. demo-felhom self-updated 0.226.1 -> 0.227.1 in under 30 seconds and then logged: [INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-08-31 06:00 CEST A box nobody deployed to now runs today's off-site integrity check on its own schedule. That is a floor DELIVERING rather than merely recording, observed instead of assumed -- and it is the strongest evidence R-242 has carried. Token hygiene: file->file, read inside the VM by a runner script, never on a command line (systemctl show ... grep -c -F token = 0). The leak grep on the committed log was PROVEN TO WORK before its 0 was believed. Teardown: guest 9100 destroyed --purge, secrets shredded AFTER the log was copied out, VM powered off, disk reverted to virgin. R-242 now records the cadence as MEASURED: five convictions and two full bakes in one day. Every bypass declared, every debt paid -- and the pattern the row exists to name is exactly that a release and its delivery are separate acts. Its other half stays open: nothing gates the VOUCH itself. |
||
|
|
99af997ab9 |
R-359/R-397 closed, R-398 corrected, R-399/R-400 filed with measured numbers
gates / gates (push) Failing after 18s
THE MEASUREMENT IS THE STORY, and it re-frames the row it was filed under. A pack was corrupted WITHOUT changing its size; plain `restic check` -- the depth that ships ON -- returned `no errors were found`, exit 0. Only --read-data caught it. So the check that shipped verifies the index, the pack inventory and the snapshot graph, and does NOT re-hash pack contents. R-399 was filed as a bandwidth-and-cadence question; it is more than that, and its row now says so. R-399 gets three MEASURED numbers instead of estimates: store 140 829 678 B / 2651 blobs / 67 snapshots; structure check 35.0 s; curve 10% 35.9 s, 50% 37.3 s, 100% 39.2 s. At this size re-reading everything costs four seconds more than reading none, because the wall clock is SFTP round-trips not transfer. The row states the limit too: these do NOT extrapolate. R-400: the sweep the task asked for found EIGHT dead debug buttons, not one. 24 endpoints referenced in debug.html, 17 dispatched. Single dispatcher, exact match, default NotFound -- so they 404. A third of a debug page does nothing, on the surface an operator reaches for when something is already wrong. R-398 is CORRECTED AND LEFT OPEN, not closed. I filed it yesterday saying resticStep is not a seam so no test can drive a restic path. The layer below it has been injectable since the off-site tier shipped. The row survives as the record that the seam EXISTS so nobody re-files it. 07 gap register: R-359 and R-397 closed; R-87 restated IN PLACE as "AND IT IS NOT R-359" because the two rows are adjacent and a check is not a restore-test. 08 alarm ladder: both event types recorded, including that `ok` is `info` and therefore mails nobody BY DESIGN, and that all three registers were checked and deliberately left alone. 00 capability map: PROVEN-LIVE for the check, the notifier and the hazard control; the scheduled firing is IMPLEMENTED only, because a week has not passed. wire_contract_gate: `offsite.last_integrity_ok` allowlisted WITH A REASON. The gate was right -- the controller emits a field no hub struct can decode. Building the display is a hub change and R-331 ruled that class the operator's decision; the entry says to delete it when a surface exists. This push used `git push --no-verify`. golden-currency is CONVICTED and right: 0.227.1 is released and the golden carries 0.226.1. A BYPASS, not a waiver, and the task spec directs it -- golden and fleet delivery are Viktor's (R-242). It is item 3 under "Waiting on you". Register 163 -> 165 -> 163. |
||
|
|
4f875174fe |
Golden 0.226.1 baked, vouched, and the fleet floor raised — the debt is paid
gates / gates (push) Successful in 17s
golden_currency_gate.py had been CONVICTED four times today across three controller releases. One bake covers all three, and the gate went red -> green on the same command, which is its proof that it measures something real. THE THREE DECLARED BYPASSES ARE NOW HISTORICAL RATHER THAN STANDING. GOLDEN_VERSION 0.226.1 GOLDEN_SHA256 70ed8e9377dec22a9b493e55f222b0e25a49d7f3caec8c506e0412fd6baefe69 size 657 197 592 B baked gitea.dooplex.hu/admin/felhom-controller:0.226.1 MinAgent 0.129.0 (read from the controller CHANGELOG header, not assumed) THE EVIDENCE IS THE ROUND TRIP, NOT THE BUILD LOG. The published bytes were downloaded back -- size and sha256 both identical to what the bake reported -- and ./etc/felhom-controller-image was read OUT of the downloaded archive: `felhom-controller:0.226.1`. That is the delivered artifact naming the controller it will start, from the bytes a customer's box would actually fetch. Acceptance markers counted, not eyeballed, each string captured from this run's own log rather than paraphrased from the runbook (two of the three the runbook named until R-233 could not match anything the script prints): docker OK (overlay2 = 1, including mount point rootfs = 1, mp0 = 1, upload OK (HTTP 201) = 1; excluding = 0, FATAL = 0, mp1 = 0. The 404 pre-gate passed before the run, so nothing was overwritten. THE VOUCH IS A THREE-FIELD CHANGE AND ALL THREE WERE CHECKED: agent_version 0.130.0 >= min_agent 0.129.0, so NOT the R-216 shape; wrapper_sha256 carried through explicitly because the handler clears it when omitted. Verified by RE-READING the manifest rather than trusting the flash -- golden option 0.226.1 SELECTED, all four shas matching. THE FLOOR IS PROVEN ACTING, NOT MERELY SET. demo-felhom self-updated within 30 seconds: "[selfupdate] Post-update startup: update successful (0.225.0 -> 0.226.1)". Both demo machines now run 0.226.1 and only one of them was deployed to by hand. Token hygiene: copied file->file, read inside the VM by a runner script, never on a command line (systemctl show ... | grep -c -F token = 0). THE LEAK GREP ON THE COMMITTED LOG WAS PROVEN TO WORK BEFORE ITS 0 WAS BELIEVED -- a throwaway copy with the token appended grepped 1, was shredded, and only then was the real log's 0 taken as evidence. Teardown: build guest 9100 destroyed --purge, secrets shredded AFTER the log was copied out (standing rule 5), VM powered off, disk reverted to virgin. R-242's OTHER half is untouched and still open: nothing gates the VOUCH itself. |
||
|
|
c8100aad6b |
Housekeeping + R-397/R-398 filed
gates / gates (push) Failing after 17s
Compresses the six rows closed today into CLOSED-ITEMS, keeping title, shipping
version, evidence paths and every sentence that states a RULE. Full original:
`git show
|
||
|
|
e027b5d999 |
Register + architecture for controller v0.226.0 (R-353/357/358/360/396), and R-395 fixed
gates / gates (push) Failing after 17s
Closes R-353, R-357, R-358 and R-360 with their shipping version and evidence path, and files two new rows. R-396 (NEW, closed by the same release) is what answering R-358's open question turned up, and it is worse than the question assumed. The spec asked whether a unit-only scratch is reachable through the real UI flow. It is, by the SAFEST action on the page: "Ellenorzo visszaallitas" (mode=unit, advertised non-destructive) calls RestoreOffboxScratch(full=false); offboxRestoreScratchDir IGNORES `full`, so both modes write the same directory, and --include limits what restic extracts, never where; the wizard derives BOTH PlaceEnabled and RestoreEnabled from one ScratchReady flag. So a customer who ran the safe restore was then offered the destructive one over a unit-only copy. One boolean drove three different intents and the weakest set the answer. R-395 (filed by the spec) is fixed in this commit, not just recorded. STATUS.md said golden 0.223.0 / floor 0.222.0 in one block and demo-hp 0.219.0 / floor 0.218.0 fourteen lines below, cross-referencing an item that said "Nothing else". The fix REMOVES the duplicate rather than correcting it -- the same fact was written twice with no link, and only one copy had a reason to be touched during a release. "What works" now points at the item above instead of restating a version. 07-backup-architecture: four rows added to the 10.2 gap register plus R-396. Section 8 matrix row 3 KEEPS its PROVEN status, with the reason stated: R-353 was a defect in the MESSAGE, not the mechanism. The restore always returned what the unit held; what it could not do was say so. A status that measures whether data comes back must not move because a status line was wrong. 00-capability-map: one new row, and it splits what is claimed. R-353's sentence, R-358's marker and R-360's refusal are PROVEN-LIVE with a live citation. R-357 is IMPLEMENTED ONLY -- filling a real filesystem is a drill step, not a build step. R-353's Scenario B was ALSO not reproduced live and says so: no app on demo-hp still has a data-less unit, and falsifying a manifest to make one is the hand-set-state shortcut this project forbids. This push used `git push --no-verify`. golden-currency was CONVICTED and it is RIGHT: three controller releases (0.224.0, 0.225.0, 0.226.0) and the golden still carries 0.223.0. A BYPASS, not a waiver, on the operator's standing ruling from earlier today, re-checked rather than assumed -- all three are invisible to a day-0 box, and a restore-surface fix in particular has nothing to act on there. The ground expires the moment a release changes first-boot behaviour. Tracked on R-242; ONE bake carrying 0.226.0 covers all three. |
||
|
|
ac6ac037bc |
R-331 live: the Backup card now reads 67 snapshots, not 0
gates / gates (push) Failing after 17s
Hub 0.109.0 + controller 0.225.0 deployed and verified from the live objects, not from a rollout message (an ArgoCD "rolled out" can name the old image): argocd sync=Synced rev==HEAD, deploy and pod both on felhom-hub:0.109.0, both boxes on felhom-controller:0.225.0 (healthy). Fetched from the live hub at the exact URL the operator's browser requests: demo-hp 67 snapshots / 134.3 MB / last success 15h ago / 50 GB quota demo-felhom 10 snapshots / 132.5 KB / last success 15h ago / 50 GB quota Both read "Snapshots 0 / Repo Size 0 MB / Integrity Unknown" before this change. Integrity row grep count is 0 on both pages. Cross-checked against the SOURCE rather than against the card itself: the boxes' own settings.json hold snapshot_count 67 / 10 and repo_size_bytes 140829678 / 135635, and 135635/1024 = 132.5 KB, matching the rendered value. Stated rather than implied: the "never measured" branch was NOT verified live. Both boxes report stats_known:true, so exercising it would have meant falsifying a box's state. It is covered at render level by TestBackupCard_ThreeWayRuling and TestBackupCard_OldControllerDegradesToUnknownNotEmpty. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM |
||
|
|
36f8630020 |
R-341 check taken, and the golden-currency bypass declared
gates / gates (push) Failing after 16s
Two gates blocked the R-331 hub push. One is FIXED, one is BYPASSED, and the difference is stated rather than blurred. FIXED -- due-checks (R-341, 5 days overdue). The +7d measurement was TAKEN on ep0 rather than deferred again. Precondition passed: proxy still MainPID 551655, ps -o lstart= still 2026-08-18 09:51:04, NRestarts=0, so this is the same proxy generation as t0 (anchor is ps, not ActiveEnterTimestamp, which reads 03:54:54Z here -- R-346's trap). Result: fd = 17. Not 17 more -- seventeen TOTAL, exactly the documented baseline, against 405 at the first check. Socket histogram: one LISTEN, ESTAB 0, CLOSE-WAIT 0. The verdict is UNANSWERABLE, not "the upgrade fixed it". R-341 asks whether the PBS 4.2.5-1 upgrade changed the fd slope; inside this interval we removed the leak OURSELVES (R-344, agent 0.130.0, now live on both boxes). A slope of ~0 measures our fix, not the upgrade, and reading it the other way would credit a changelog that was read in advance and found to contain no such mechanism. The perturbation pre-registered for this window was Phase C at ~3%; the actual perturbation was the removal of the entire phenomenon. Row closed as moot. What it DOES establish is worth more than the original question: twelve days after the R-344 fix, same proxy generation, no restart to hide behind, ep0 sits at baseline with zero established connections. R-336's ~323-day runway concern retires with it. BYPASSED -- golden-currency. Controller v0.224.0 and v0.225.0 are released and the newest golden bake carries 0.223.0, so a machine installed right now gets neither. The gate is RIGHT. This push therefore uses `git push --no-verify`, declared here, in hub/CHANGELOG.md, in REPORT.md and on R-242. A BYPASS, not a waiver: the gate offers a waiver only for a release that DELIBERATELY needs no golden, and these need one. The operator was asked and ruled bypass-now-bake-later, on the ground that neither fix bites a day-0 box -- R-330 is a nightly false alarm about apps a new box has not installed yet, R-331 is a hub display over backups a new box has not taken yet -- and both arrive by self-update. That ground is recorded because it is what to re-check: it does NOT extend to a release changing first-boot behaviour. OWED: bake a golden carrying 0.225.0 and vouch it (RUNBOOK-manual-build.md 4.1, three-field change, MinAgent 0.129.0). Fourth bypass of this gate, and the gap is now two releases wide rather than one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM |
||
|
|
f5c9411e5e |
R-331 (hub half): the Backup card reads offsite, not the dead backup fields (v0.109.0)
The customer page's Backup card read `Snapshots 0 / Repo Size 0 MB / Integrity
Unknown` for EVERY customer, indefinitely. Measured on demo-hp 2026-08-30 while
that night's controller log said `[offbox] backup OK: 8 app(s) backed up, 67
snapshot(s), 2m14s` and the box held snapshot_count:67, repo_size_bytes:
140829678, stats_known:true.
A card reading "no backups" over a working backup is worse than no card -- the
R-88 direction of failure (degrade to NO BACKUP rather than to UNKNOWN) on the
one screen that answers "is this customer protected?".
The data was never missing. The card rendered the report's `backup` object,
whose snapshot/size/integrity fields have had no producer since slice 8C. The
live numbers are in the `offsite` object, which THIS PACKAGE already reads for
the Offsite page and which monitor.OffsiteChecker already alarms from. Proof the
bytes were arriving: the Offsite page rendered demo-hp's usage as 0.1 GB from
that very object while the Backup card said 0 MB. So this is a render fix over
an existing feed, not a new pipeline.
Not a one-line swap, because snapshot_count:0 means two opposite things --
"holds nothing" and "never measured". R-225 measured that confusion one layer
down. backup_card.go resolves a three-way ruling in Go (a {{if}} chain over
map[string]interface{} float64s cannot keep the absent/zero distinction the card
is entirely about):
no offsite object -> "No off-site data reported", and says explicitly that
this is NOT the same as "no backups"
disabled + state -> names the blocker (needs_credential)
stats_known:false -> em-dash + "never been measured". NEVER 0
stats_known:true -> the real numbers, INCLUDING a real 0
A pre-v0.225.0 controller sends no stats_known -> false -> "unknown". That
direction is pinned: upgrading the hub ahead of the fleet must not report every
un-upgraded customer as having zero backups.
The Integrity row is DELETED, not re-sourced: nothing produces it, the
controller runs no integrity check, and NotifyIntegrityOK/Failed are called from
nowhere.
RED-PROOF: restore the pre-fix card markup -> all four tests fail, reporting 67
and 134.3 MB absent from the rendered page and the Integrity row present. The
tests drive handleCustomerUnified and grep the HTML on purpose: the defect was
the template's choice of source object, so a test one layer below it would have
been green against the shipped bug.
Green gate clean: 18 packages, rc 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
|
||
|
|
c2c1fb48dc |
REPORT: record the pushed commit hash
gates / gates (push) Failing after 17s
|
||
|
|
c30430c530 |
skills: five process-domain skills + check_skills.py
gates / gates (push) Failing after 15s
The four existing skills cover the product; nothing covered how work is reported. Two rules this project has paid for — check the artifact rather than the report, and do not state a claim more firmly than the evidence allows — lived only in the operator's head and in chat, where Claude Code never read them. - felhom-evidence five confidence tiers, artifact-over-report - felhom-diagnosis no hypothesis until a command has been seen red - felhom-plain-language ASD-STE100, two options, the re-pitch - felhom-handoff the note goes to a FILE, not the conversation - felhom-doc-authoring the pointer decides whether material is reached scripts/check_skills.py asserts what decides whether a skill is EVER reached: frontmatter parses, name == directory, description and body non-empty, under 150 lines, installed copy still samefile()s into the repo. install_skills.py globs and never reads the file, so a missing description installs perfectly and then silently never loads. It convicted on its first run: felhom-build-deploy is 179 lines. NOT trimmed here (pre-existing skills are out of scope, and trimming a deploy skill without exercising its commands is how a wrong command reaches a live host) — a named single-entry GRANDFATHERED exception, WARNed every run, R-394. A new skill over the limit is convicted. Red-proof run and seen failing: description removed from felhom-evidence -> exit 1, "frontmatter field 'description' is missing or empty". Restored, tree clean. skills/SOURCES.md records both MIT upstreams, that these are adaptations not copies, and the six pieces deliberately EXCLUDED with reasons. Register: R-392 (no architecture doc covers the two-AI workflow), R-393 (decision-log skill deferred, with the reason), R-394. |
||
|
|
ebdc04601d |
docs(hub v0.108.0): the delivery grain, the cooldown ruling, and gate 11's first subject
gates / gates (push) Successful in 15s
The alarm ladder gains §6.2 - which events are per-app, per-run, per-tier or coarse, and why the default is coarse. CONTEXT records two rulings: the grain is allow-listed rather than inferred from the payload, with crossdrive_failed as the proof that a payload rule would have been wrong; and a finding recorded only in REPORT.md has a lifetime of one session. R-389 closed and compressed, keeping its rules and naming the commit whose git show returns the full text. R-390 and R-391 left open. REPORT.md is gate 11's first real subject and passes: six observations, two FILED, four NOT-A-FINDING with their reasons. Three of those declarations are things a tidier report would have omitted - the gate's own spec would have passed the item it was built to catch, the burst has no ceiling, and ArgoCD said "successfully rolled out" while still running the old image. STATUS carries forward the one thing outstanding: the controller floor still reads 0.222.0 while the golden reads 0.223.0. |
||
|
|
45659bdc5a |
hub v0.108.0: deploy the per-app cooldown grain (R-389)
gates / gates (push) Successful in 17s
Image built and pushed to the registry BEFORE this manifest bump lands, so a sync can never point at a missing tag. Auto-sync is off; the sync that follows is deliberate. Never kubectl set image. |
||
|
|
2fc4a15fa3 |
R-389: key the operator cooldown per app for app_start_failed; gate 11 makes an unfiled observation refuse the push
gates / gates (push) Successful in 16s
The cooldown key was customerID:eventType plus the tier and run suffixes, and none of them names an app, so every app going down inside the same hour collapsed onto one key and only the first was mailed. Measured on demo-hp: bookstack sent 09:27:51, privatebin suppressed 09:31:51 under key=demo-hp:app_start_failed. cooldownStackSuffix is the third sibling of cooldownTierSuffix and cooldownRunSuffix, and separate for the reason the second one's docstring already gives: the existing two keep byte-identical semantics for every type that uses them. It is ALLOW-LISTED to app_start_failed and takes the event type as well as the details, unlike its siblings, and that asymmetry is the safety property. The backup family's cooldown is coarse ON PURPOSE (R-97a, R-182) so one full disk sends one digest rather than one mail per app - and crossdrive_failed is severity error, reaches the operator leg, and carries stack_name through a DIFFERENT struct, so a payload-shape rule would have split it silently. The hour itself does not change. Gate 11 refuses a push whose REPORT.md carries an observation with neither `FILED: R-NNN` nor `NOT-A-FINDING: <reason>`. It deliberately does NOT accept a passing mention of some other R-number: the lost item cited R-182 as an analogy, so "cites a register row" would have passed the very item the gate exists to catch. That discrepancy with the spec is recorded in the gate's docstring. Registered here and in the controller and agent runners. NOT in the catalog runner - it has no shared-gate mechanism and appends --all to every gate; filed as R-391 rather than left as a sentence, which is this session's lesson. PROMPT-TEMPLATE.md §15.9 corrected: "documented, NOT acted on" was the wording that invited the gap, and it now names the markers and points at the gate. R-390 filed for the golden-bake runbook's missing `pveam update`. Hub tests 709 -> 716. |
||
|
|
f751aea4e3 |
R-389: file the cooldown-grain finding that was never filed
gates / gates (push) Successful in 16s
Only the first broken app per hour reaches the operator. The cooldown key is customerID:eventType plus the tier and run suffixes, and neither reads an app name, so every app that goes down inside the same hour collapses onto one key. Measured on demo-hp 2026-08-23: bookstack sent at 09:27:51, privatebin four minutes later logged `suppressed - operator cooldown 1h, key=demo-hp:app_start_failed`. Filed FIRST, before any code, for two reasons. It should have existed since yesterday and did not - it lived in a REPORT.md observations paragraph and nowhere else, which is R-341's shape one surface over. And the gate this session adds refuses a push whose report carries an observation with no row behind it, so the row has to precede the gate or the gate refuses its own commit. |
||
|
|
2f7c9a6ce5 |
docs(R-329/R-386/R-387): the severity contract, the intent ruling, and Part 5 recorded
gates / gates (push) Successful in 17s
The alarm ladder gains the severity contract (the hub's vocabulary is exact, it coerces silently, and three things now hold it) and the intent test with its three-way ruling on unknown. Both marked [DESIGN] with the live measurements. Part 5 is RECORDED AND NOT IMPLEMENTED: the operator's notification philosophy, verbatim, marked plainly as direction rather than current behaviour, with the 12 -> 15 toggle growth as the argument. Filed as R-388, a product decision. R-329 and R-386 compressed into CLOSED-ITEMS with their rules kept and the full-text commit named. R-387 filed closed - including WHY the dispatcher branch was kept rather than deleted, which is evidence (three monitor checkers call ProcessEvent directly) and not caution. The drill record names three things that had to be re-run: an inert red-proof mutation, Scenario G refused twice behind an HTTP 200, and the live Scenario A NOT proving the customer gate because demo-hp has no prefs row at all. Register: OPEN 328325 -> 328132 B, CLOSED 71441 -> 74642 B. |