Baked now, off the weekly cadence, because check-family-gate.py refuses a family_gate template while the newest
baked golden is older than 0.287.0. Round trip sha == bake; three-field vouch; R-120 refused 0.286.1 after.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- Decision 52: catalog 6a3ead9 (re-test entries, gates, decoys, the monthly command); proven end to end on 9202
through the leg; runbook monthly-floating-retest.md; nothing to re-test on the engine lines today.
- Decision 53 + R-741: controller v0.284.2 (0.284.0/0.284.1 never floored — two wiring faults found live on 9202);
floor 0.284.2; one-time sweep 9202 26.6 -> 5.7 GB, demo-hp 24.3 -> 13.5 GB; the install hold proven as a stranger.
- Golden 0.284.2 baked, round-trip identical, vouched (agent 0.138.0, min_agent 0.131.0); the gate prints OK.
- Rows 377 -> 383: opened R-743..R-748, closed R-736, R-737, R-740, R-741, R-748; narrowed R-739, R-698, R-446.
- register_shape_gate: a lettered id (R-88a) is a row too (R-748), with a decoy seen red.
Evidence: documentation/audits/night-rulings-2026-09-30/, documentation/tests/golden-0.284.2-2026-09-30/.
Report: REPORT-night-rulings-2026-09-30.md.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The volunteer guide has an English twin. It is a TRANSLATION, not a rewrite: 16
sections in the same order, identical step counts, table rows and warning blocks per
section (measured, 0 sections differing in structure). Word counts are NOT a twin —
English runs 19 % longer overall and up to 42 % on the short sections, because
Hungarian is agglutinative; the +-15 % criterion the task asked for does not survive
contact with this language pair, so structure is the measure reported instead.
Golden 0.258.0 baked, published and vouched, with its record. One run, no aborted
attempts: the 0.246.0 bake's two traps were both avoided by following its own record.
Token proven not to leak with a planted control before the zero was believed.
The waiver is NOT retired, and the record says why in one line: it is the mechanism of
operator ruling R-468, not a note about this golden, and deleting it would turn the
next release without a bake red immediately. It is also not load-bearing today.
R-595: the catalog's copy gate could not run in CI at all — six pushes red, six alarm
mails, while the local hook was green. Found by reading the operator's inbox, not by
anything in the session that caused it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Live at iso.felhom.eu, sha256 dceacae5da247d76cad065bf6c0d3bbefac8d8a5f8e571
db2fd8449a97e94829, and both download pages now name it.
Every gate criterion is recorded with its OBSERVED value in
documentation/tests/iso-release-1.29.0-2026-09-18/ — including two proof
installs from the published bytes, one per boot-menu entry, each with a first
boot AND one reboot: /etc/issue bilingual with zero hits for 8006, pvebanner
masked, package 1.29.0 installed, unit enabled and fired, pairing code present,
and the installed script byte-identical to repo HEAD. G11: the downloaded bytes
hash to the published checksum.
A defect was caught BETWEEN builds by looking at the screen rather than at the
config: the second menu entry read "Felhom telepítés (szöveges mód) / Install
Felhom (text mode)" — 58 characters — and the GRUB menu box cut it at "Instal".
The English half was unreadable on the boot screen. Shortened to "… / text" and
rebuilt; the published image is the rebuilt one. The Hungarian half is the part
that may not change, so the English half is the part that gave.
Teardown: VMs 323/324/325 destroyed, the two unclaimed appliance registrations
discarded (zero left in `registered`), guest 9201 untouched — 23 containers
before and after. The two stale *.rootpw.txt files were shredded from the
publish source directory before the upload ran from it (R-587, files gone; the
guard that would stop it recurring is still open).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Decision 1 delivered: floor 0.246.0 served, the N100 on 0.246.0 within seconds;
Peti's box is DOWN on the hub and receives it when it reports.
Decision 2: the agent vouch was refused by R-120 until a newer golden existed.
On the operator's choice, golden 0.246.0 was baked (sha 05b7559d, amd64, all
markers, token leak 0 with control 1, registry 200 before teardown) and vouched
together with agent 0.132.0. golden_currency_gate: WAIVED -> OK. Bake evidence
filed where the gate and runbook read it: tests/golden-0.246.0-2026-09-17/.
Recorded, none reaching the registry: a first attempt on the arm64 template
(my version sort), a self-matching pkill, and an OOM-killed watcher whose
post-bake steps were done by hand.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
ISO 1.27.1 gated PASS and proven live: first-boot console Felhom-only,
pvebanner masked across a proven reboot. Hub v0.113.0 hand-over copy live
(R-497 closed). Full first hour walked again on customer tester-1 (three
disks + one disk): deploy, use, backup, removal, byte-identical restore,
power cut, typo all PASS. The tunnel gives a fresh box no routes: 12/12
503 from DooPlex (R-505); the record has no e-mail (R-508). Rows R-507,
R-508 filed; R-496/R-495 fixed/answered awaiting publish; day-0 A.1 no
longer claims the controller creates hostnames (R-506). NOT PUBLISHED.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Golden 0.242.0 baked, round-trip verified and vouched (cadence rule, R-468).
Fresh box from the public ISO on demo-hp: landed on the vouched set, two apps
deployed and used, backup, remove, byte-identical restore, power cut and code
typo all PASS. Stopped for a volunteer by R-493 (no instructions) and R-494
(the setup mail's dashboard link has no DNS; intervention I1). R-493..R-500
filed. Capability map: first-hour row added (PARTIAL), journey row scoped.
Stopgap Hungarian volunteer guide written. Hub teardown layer pending.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Operator rulings 2026-09-13, both shipped the same day:
- MariaDB finishes its own conversion (catalog eec1228/bd32830/3525e35). Harness E3/E3b `proven`
with engine_state_after "already upgraded to 12.3.3-MariaDB [exit=1]", the skip line gone, C3
still `failed`; landed on demo-hp through the real 15-min cycle, nothing recreated, one deliberate
restart logged "MariaDB upgrade not required" with the app serving. Evidence:
documentation/audits/r459-close-2026-09-13/. The engine-major rule + gate keep every engine
inside its major until Slice 4 (R-448) — removal tracked as R-469.
- Goldens on a cadence, not per release. golden_currency_gate.py reads a dated waiver
(documentation/tests/golden-waiver.yml, <= 14 days, row-bound): valid + BEHIND -> loud advisory,
exit 0; expired -> red again naming the date; UNRECORDED (R-385) never covered; malformed -> 2,
never 0. Tests cases 5-15 incl. the R-421 decoy; red-proof old-vs-new on the real behind tree.
R-242's vouch half stays open. Cadence in RUNBOOK-manual-build.md §4.2 + the checklist.
- Golden 0.236.0 baked, round-tripped, vouched, floor raised 0.232.0 -> 0.236.0
(documentation/tests/golden-0.236.0-2026-09-13/) — the last per-release bake; the waiver was
issued AFTER it landed. No --no-verify anywhere in this session.
Rows: R-459 CLOSED, R-467 CLOSED, R-242 narrowed; R-468/R-469/R-470/R-471 opened. 09 §3 gains
decisions 5 and 6; STATUS items 11 and 12 closed; CONTEXT records the cadence ruling.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
09-update-architecture.md gains the fourth dated operator ruling (2026-09-06,
Option 1) and its section 5 is rewritten from a proposed shape into the shipped
one: the pin, the stored definition, the render table, the four writers, the
startup ordering, and the trap this slice set for slice 2 - the live compose file
is now the frozen one, so a badge comparing against it would answer Naprakesz on
exactly the apps that are behind.
02-controller-module-map.md said 'copy compose + .felhom.yml'. That stopped being
true today, so it is corrected, and the two sections describing the old seam now
carry a banner saying they describe v0.234.0 and below - kept because every box
under v0.235.0 still behaves that way and because they are the measured account
of why it changed.
R-447, R-441, R-438 and R-455 closed and compressed into CLOSED-ITEMS; R-458
opened for the .felhom.yml asymmetry, with what would settle it by measurement.
Live evidence: two real catalog pushes travelling the real 15-minute cycle, both
reverted, the tree byte-identical afterwards. The restart that used to take 18.3
seconds and pull a new image now takes 0.1 seconds and pulls nothing.
The operator looked at demo-felhom and found OpenGist - up 15 hours, running
exactly the catalog pin, showing no badge at all. 09-update-architecture.md had
recorded that as an accepted limitation the day before: 'the fleet view fills in
gradually'. On a quiet box gradually means never, and a feature that fills itself
in on an event nobody triggers is, on the quiet installations, not shipped. That
limitation row is now struck with the reason kept.
The living document gains slice 1b, the two admission rules of the backfill (it
never overwrites, and it refuses to seed a partial observation because the badge
reads a service-count mismatch as BEHIND), and the note that the same field having
two writers with two different admission rules is deliberate.
Live evidence added: all nine apps already had records by the time 0.234.0 was
ready, so the natural fleet state could no longer exercise the new code - said
plainly rather than papered over. The pre-0.233.0 shape was recreated on demo-hp
by stripping two records; the backfill re-seeded exactly those two with digests
matching independently-read ground truth and left the other seven alone.
The refusal half was deliberately NOT staged live: it needs a degraded app, and
manufacturing one risks the false-customer-email class that already cost 61 mails
(R-330). Unit-tested with a red-proof, and recorded as unproven-live.
R-457: a test that hardcodes a date and asserts an age derived from it is green
only on the day it is written. Mine was, and it went red overnight. Six other
files carry both a date literal and time.Now() - named as candidates, not accused.
I reported that the vaulted dashboard password no longer worked on either demo
box, and quoted the controller's own 'Failed login' as the discriminator. The
password was fine. ~/.config/credentials quotes its values with SINGLE quotes and
my sed stripped only double quotes, so the quote characters went out as part of
the password. The operator corrected it in one line; one retry returned 302.
The instrumentation lesson is the finding and R-453 now carries it: 'Failed login'
separates wrong-password from wrong-Host-header, and that is ALL it separates. It
cannot tell a wrong password from wrong password HANDLING, and I read it as if it
could. This is the second time this file's quoting has produced a confident wrong
verdict, so the fix is one shared extraction helper, not a resolution to be careful.
With the session recovered, the badge is validated on live pages: Naprakesz twice
on /stacks and on /apps/bookstack; NO badge at all on /apps/docmost (a deployed app
with no record - absent is UNKNOWN, not current); and 'Frissites elerheto - 52
napja' on both surfaces, the age being real arithmetic on bentopdf's catalog_since.
The behind state was staged by editing one compose tag, with no restart and no
up -d, and reverted byte-identically (sha256 equal, diff empty, container never
touched). Capability-map row upgraded to PROVEN-LIVE with the one unexercised
badge state named. STATUS item 9 now needs nothing from the operator.
R-438's document half. It records how an update works AS MEASURED, quotes the
RestartStack comment that proves the restart half was CHOSEN (a design decision
is not a defect), carries the three operator rulings of 2026-09-02, strikes the
word 'rollback' (once a migration has run the old image will not start), states
the target shape, and lists the seven slices with a status each.
R-438 and R-440 amended and BOTH STAY OPEN: the mechanism is documented, not
changed. Nothing closed, so CLOSED-ITEMS.md is untouched.
Eight new register rows, 194 -> 202: R-446 (Naprakesz can be false for the 23
floating pins), R-447..R-451 (one per remaining slice, with a rank and an owner),
R-452 (no gate enforces catalog_since - the runner fetches at --depth 1), and
R-453 (the vaulted dashboard password is stale on BOTH demo boxes, which is what
stopped the badge render from being validated live).
Live evidence for slices 1 and 2 in documentation/tests/. The record is PROVEN
LIVE through the boot reconciler on demo-hp - one entry per compose service,
digests matching ground truth read independently. The badge RENDER is not, and
the five attempts are listed rather than summarised.
GOLDEN_SHA256 5f8a53ed5b19a6cb2006298ce6239f6fca2b990cc3ef6eada89f602801ca91b8,
657 494 489 B. 0.231.0 was never baked, so the fleet went 0.230.0 -> 0.232.0.
THE CHECK THE 0.230.0 BAKE SKIPPED, AND THIS ONE DID NOT: the bake script's fingerprint was
compared ACROSS THE HOP - 7b0fb5cf...73b6a1 on DooPlex and inside the VM. The previous bake
recorded only the DooPlex-side hash and said so; this one is a measurement.
Three independent readers agreed before anything was vouched: the bake's own print, the round
trip of the published bytes (HTTP 200, 657494489 B, same sha, hashed from what was
downloaded), and the hub's Day-0 dropdown reading Gitea on a different code path. The
delivered artifact names its own controller - ./etc/felhom-controller-image reads
felhom-controller:0.232.0 - with 19382 entries under var/lib/felhom/docker/.
Both pre-gates were shown able to see something before their zeroes were believed, and the
manifest was RE-READ after vouching rather than trusted from the 303 flash.
DELIVERY WAS ACTUALLY EXERCISED. Both boxes had been hand-deployed during validation, so the
floor had nothing to move. Rather than report delivery untested, demo-felhom was rolled back
to 0.231.0 and the chain run for real - it moved itself in ~20s:
10:47:16 controller-swap: image file written, restarting bootstrap target=...0.232.0
10:47:26 controller-swap: new controller healthy target=...0.232.0
And this is the first bake golden_currency_gate.py actually gates: it now reads the
GOLDEN_SHA256 line out of the bake log rather than matching a directory name (R-410, shipped
hours earlier the same day). All 13 felhom.eu gates are green, golden-currency included, for
the first time since v0.230.0 was released.
Controller v0.231.0 + hub v0.110.0, both deployed and verified on demo-hp.
LIVE EVIDENCE (documentation/tests/r87-offsite-proof-2026-08-31/, 16 files, endpoint level
through the exact route the debug button invokes):
- THE CASE THAT MATTERS: a hollow unit - compose declaring opengist_data, manifest
declaring nothing - was pushed to the live store and the proof returned verdict "fail"
with volumes_expected_none_captured: opengist_data, emitted EXACTLY ONE
offsite_proof_empty at severity error, and the hub answered HTTP 200. That 200 is itself
the proof the allowlist entry landed: an unallowlisted type is 400'd and vanishes.
- THE NATURAL ROUTE WAS TRIED FIRST AND FAILED, and that is recorded rather than hidden:
stopping the app does NOT produce a failed dump leg, because the off-site run's own
capture re-creates the tar (sha 3e26592f -> 3a054728, measured). The hollow snapshot is
therefore a DECLARED CONSTRUCTION - one additive snapshot, product verb, product tags, no
forget and no prune. State restored: the product's own run made a healthy snapshot the
newest again and the proof then passed opengist.
- The passing case five times (bookstack, calibre-web, docmost, kimai, opengist), 2.2-4.0s
each, matching the spike's measured band.
- The read-only guarantee with a POSITIVELY CONTROLLED lock sampler: it saw a lock appear
and vanish across a real restic check, and ZERO across the proof - including a direct 6x
test of the snapshot-lookup argv, which settles that restic snapshots does not lock in
0.14.0 either.
- Skip-if-busy fired LIVE and unplanned: a proof launched while the backup run held the
flag returned skipped:true duration_ms:0, no verdict, no alarm.
- The customer's own verification copies were untouched throughout, which is the safety
property the separate proof root exists for.
ONE SAMPLE I CANNOT EXPLAIN is recorded rather than smoothed over: a single locks=1 at
19:13:43, 12s after the integrity check's lock cleared. Two independent tests exclude the
proof; I did not establish what it was.
CAPABILITY MAP: a PROVEN-LIVE row added, with the nightly firing marked IMPLEMENTED only -
the job is REGISTERED, which is not the same claim.
07 section 8 MATRIX ROW 4 WAS NOT MOVED, deliberately, and section 10.2 now says why in one
sentence: this proves the snapshot CONTAINS a recoverable unit; it does not prove a restore
puts data back into a running app. Without that sentence the new green tick reads as
covering the drill.
REGISTER: R-87 CLOSED and compressed into CLOSED-ITEMS.md. OPEN 172 -> 171, CLOSED 151 ->
152. No new rows minted. R-408 and R-409 stay open and are referenced by this work.
golden-currency is RED and it is a DECLARED, EXPECTED debt: v0.231.0 is released and the
newest golden carries 0.230.0. The fleet is on 0.230.0; demo-felhom does not have this job.
A golden carrying 0.231.0 is OWED and it is Viktor's call (R-242). This push uses
--no-verify for that reason - bypass #8.
GOLDEN_SHA256 9287f7cef5f13166276e8406005e3f28004004510c5184f1c1c7377f7aafad2e,
657 873 700 B. Evidence documentation/tests/golden-0.230.0-2026-08-31/.
WHY IT WAS OWED: the newest golden was 0.229.0, which IS the build R-403 says deletes a
good copy. Every fresh install and the whole fleet floor still carried it.
golden_currency_gate.py had been red across dddcc80, 6e550ae, 130f7a6 and 32a4c35.
THREE INDEPENDENT READERS agreed before anything was vouched: the bake's own print, the
round trip of the PUBLISHED bytes (HTTP 200, 657873700 B, same sha), and the hub's Day-0
dropdown reading Gitea on a different code path. And the delivered artifact names the
controller it will start - ./etc/felhom-controller-image read OUT of the downloaded
archive says felhom-controller:0.230.0, with 19382 entries under var/lib/felhom/docker/.
BOTH PRE-GATES were shown able to see something before their zeroes were believed: the
404 pre-gate, and the token-leak grep which returns 0 on the committed log and 1 on a
seeded throwaway copy. The transient unit's own properties were grepped for the token
too - 0, with the same seeded positive control returning 1. Acceptance markers counted on
the COMMITTED log: 1/1/1/1 present, 0/0 absent, and the zeroes are believable because the
same including-mount-point pattern returns two real lines on that file.
THE VOUCH IS A THREE-FIELD CHANGE and only one field moved, which is stated rather than
left to look careless: golden_version 0.229.0 -> 0.230.0; agent_version 0.130.0 and
min_agent 0.129.0 UNCHANGED because v0.230.0's CHANGELOG header says MinAgent 0.129.0 and
0.129.0 <= 0.130.0, so this is not the R-216 shape. The 303 flash was not treated as
proof - the page was re-read and golden_behind_fleet confirmed absent.
THE FLOOR is a separate setting and was raised on the operator's explicit answer:
min_controller_version 0.229.0 -> 0.230.0. THE POSITIVE OBSERVABLE, from the agent's own
journal on demo-felhom, which was still running the defective 0.229.0:
16:21:30 controller-swap: image file written, restarting bootstrap target=...0.230.0
16:21:40 controller-swap: new controller healthy target=...0.230.0
Both boxes now 0.230.0 healthy. Honest note: the polling loop's first read already said
0.230.0, so the transition was not seen by the loop - the journal is the evidence.
R-410 FILED, found while the gate went green: golden_currency_gate.py is satisfied by a
DIRECTORY NAME (EVIDENCE_RE against os.listdir, :89,:123). I created the evidence
directory before the bake finished and the gate would have passed at that moment. It
already declares that it does not check the vouch; it does not declare that the bake
check is a filename check. Fix: read the GOLDEN_SHA256= line out of the directory's
bake.log, with a red-proof on an empty directory.
R-242 updated - seventh debt, paid the same day, twice in one day.
Teardown: pct destroy 9100 --purge, shred -u AFTER the log was copied out, poweroff,
qemu confirmed exited with ps -eo comm (not pgrep -f, which self-matches), disk reverted
to virgin.
All 13 gates green - the first push this session that needed no --no-verify.
Ceiling R-409 -> R-410.
GOLDEN_SHA256 39aa886df77b21757aef3b298a389343dc0df5134bb0f14e8f92a451d7bdae87, 656 864 331 B.
The evidence is the ROUND TRIP, not the build log: the published bytes were downloaded back and
match the bake on both size and sha, and ./etc/felhom-controller-image read OUT of the downloaded
archive says felhom-controller:0.229.0 - the delivered artifact naming the controller it will start.
A THIRD independent reader agreed before anything was vouched: the hub's own Day-0 dropdown read the
same sha straight from Gitea, a different code path.
Both pre-gates were proven able to see something before their negative results were believed - the
404 pre-gate against a 200 from 0.228.0, and the token-leak grep against a seeded throwaway copy.
Acceptance markers counted on the COMMITTED log: 1/1/1/1 present, 0/0/0 absent.
The vouch is a three-field change, checked rather than assumed: MinAgent 0.129.0 read from the
golden's controller CHANGELOG header, agent_version 0.130.0 >= min_agent 0.129.0 (not the R-216
shape), agent_sha256 and wrapper_sha256 carried through explicitly because the handler clears a
field it is not sent. Verified by re-reading the manifest, never by trusting the flash. The R-120
gate PASSED rather than being bypassed - fleet newest 0.229.0, golden 0.229.0.
The floor is proven ACTING, not merely set: demo-felhom self-updated 0.228.0 -> 0.229.0 and logged
settle-gate GO at/above floor 0.229.0. Nobody deployed to that box. Both demo machines now carry the
Tier-2 unit restore.
golden_currency_gate.py went red -> green; the --no-verify bypass declared on c2de785 is now
historical. R-242's other half is untouched and still open: nothing gates the VOUCH itself.
Teardown: build guest 9100 destroyed --purge, token/runner/script/log shredded AFTER the log was
copied out, VM powered off, qemu confirmed gone from ps -eo comm, disk reverted to virgin.
Bake: build-golden.sh v3.0.0 in the drill VM, reverted to virgin and cold-booted.
GOLDEN_SHA256 76a3a98b9e7cc23bf8ae51b38a6272f576df285cb34cd22235ac3f06a31e53ec,
658 079 744 B. All four acceptance markers counted 1; excluding/FATAL/mp1 counted 0.
The 404 pre-gate was proven to work before its 404 was believed — the target URL
404'd while the existing 0.227.1 package 200'd on the same command.
The evidence is the ROUND TRIP: the downloaded bytes match the bake's size and
sha, and ./etc/felhom-controller-image read OUT of the downloaded archive says
gitea.dooplex.hu/admin/felhom-controller:0.228.0.
Vouch: three fields together — golden_version 0.228.0, agent_version 0.130.0,
min_agent 0.129.0 (read from the controller CHANGELOG header, not assumed);
wrapper_sha256 carried through explicitly. agent >= min_agent, so not the R-216
shape. Verified by RE-READING the manifest, never the flash. R-120 gate passed.
Floor raised 0.227.1 -> 0.228.0 (impact preview {"below":3,"valid":true}).
The floor is ACTING: demo-felhom self-updated 0.227.1 -> 0.228.0 in under a
minute and re-registered offsite-integrity by itself. Both demo boxes now
re-read their whole off-site store on the weekly check.
Token hygiene: file->file scp, runner script inside the VM, unit properties
grepped 0. The leak grep on the committed log was proven with a planted copy
(1) before its 0 was believed. Teardown: guest 9100 purged, secrets shredded
after the log was copied out, VM off, disk reverted to virgin.
golden_currency_gate.py red -> green. All 12 felhom.eu gates OK.
Second full delivery of the day. golden_currency_gate.py went red -> green on
the same command, so the --no-verify bypass declared on the previous push is now
historical rather than standing.
GOLDEN_VERSION 0.227.1
GOLDEN_SHA256 66754491dc9bd0130ef8ded9562f63c53a5ffdcfd91baa551141e55fa083ea32
size 657 403 203 B
baked gitea.dooplex.hu/admin/felhom-controller:0.227.1
MinAgent 0.129.0 (read from the controller CHANGELOG header, not assumed)
THE EVIDENCE IS THE ROUND TRIP. The published bytes were downloaded back -- size
and sha256 identical to what the bake reported -- and ./etc/felhom-controller-
image was read OUT of the downloaded archive: felhom-controller:0.227.1. That is
the delivered artifact naming the controller it will start, from the bytes a
customer's box would actually fetch.
Markers counted: docker OK (overlay2 = 1, mount point rootfs = 1, mp0 = 1,
upload OK (HTTP 201) = 1; excluding = 0, FATAL = 0, mp1 = 0. 404 pre-gate passed
before the run and the script's own pre-delete agreed, so nothing was
overwritten.
Three-field vouch, all three checked: agent_version 0.130.0 >= min_agent 0.129.0
(NOT the R-216 shape), wrapper_sha256 carried through explicitly because the
handler clears it when omitted. Verified by RE-READING the manifest rather than
trusting the flash. The R-120 gate on that POST passed on its own terms rather
than being worked around.
AND THE LINE WORTH KEEPING. demo-felhom self-updated 0.226.1 -> 0.227.1 in under
30 seconds and then logged:
[INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-08-31 06:00 CEST
A box nobody deployed to now runs today's off-site integrity check on its own
schedule. That is a floor DELIVERING rather than merely recording, observed
instead of assumed -- and it is the strongest evidence R-242 has carried.
Token hygiene: file->file, read inside the VM by a runner script, never on a
command line (systemctl show ... grep -c -F token = 0). The leak grep on the
committed log was PROVEN TO WORK before its 0 was believed.
Teardown: guest 9100 destroyed --purge, secrets shredded AFTER the log was
copied out, VM powered off, disk reverted to virgin.
R-242 now records the cadence as MEASURED: five convictions and two full bakes
in one day. Every bypass declared, every debt paid -- and the pattern the row
exists to name is exactly that a release and its delivery are separate acts. Its
other half stays open: nothing gates the VOUCH itself.
THE MEASUREMENT IS THE STORY, and it re-frames the row it was filed under. A
pack was corrupted WITHOUT changing its size; plain `restic check` -- the depth
that ships ON -- returned `no errors were found`, exit 0. Only --read-data
caught it. So the check that shipped verifies the index, the pack inventory and
the snapshot graph, and does NOT re-hash pack contents. R-399 was filed as a
bandwidth-and-cadence question; it is more than that, and its row now says so.
R-399 gets three MEASURED numbers instead of estimates: store 140 829 678 B /
2651 blobs / 67 snapshots; structure check 35.0 s; curve 10% 35.9 s, 50% 37.3 s,
100% 39.2 s. At this size re-reading everything costs four seconds more than
reading none, because the wall clock is SFTP round-trips not transfer. The row
states the limit too: these do NOT extrapolate.
R-400: the sweep the task asked for found EIGHT dead debug buttons, not one. 24
endpoints referenced in debug.html, 17 dispatched. Single dispatcher, exact
match, default NotFound -- so they 404. A third of a debug page does nothing, on
the surface an operator reaches for when something is already wrong.
R-398 is CORRECTED AND LEFT OPEN, not closed. I filed it yesterday saying
resticStep is not a seam so no test can drive a restic path. The layer below it
has been injectable since the off-site tier shipped. The row survives as the
record that the seam EXISTS so nobody re-files it.
07 gap register: R-359 and R-397 closed; R-87 restated IN PLACE as "AND IT IS
NOT R-359" because the two rows are adjacent and a check is not a restore-test.
08 alarm ladder: both event types recorded, including that `ok` is `info` and
therefore mails nobody BY DESIGN, and that all three registers were checked and
deliberately left alone. 00 capability map: PROVEN-LIVE for the check, the
notifier and the hazard control; the scheduled firing is IMPLEMENTED only,
because a week has not passed.
wire_contract_gate: `offsite.last_integrity_ok` allowlisted WITH A REASON. The
gate was right -- the controller emits a field no hub struct can decode.
Building the display is a hub change and R-331 ruled that class the operator's
decision; the entry says to delete it when a surface exists.
This push used `git push --no-verify`. golden-currency is CONVICTED and right:
0.227.1 is released and the golden carries 0.226.1. A BYPASS, not a waiver, and
the task spec directs it -- golden and fleet delivery are Viktor's (R-242). It
is item 3 under "Waiting on you".
Register 163 -> 165 -> 163.
golden_currency_gate.py had been CONVICTED four times today across three
controller releases. One bake covers all three, and the gate went red -> green
on the same command, which is its proof that it measures something real. THE
THREE DECLARED BYPASSES ARE NOW HISTORICAL RATHER THAN STANDING.
GOLDEN_VERSION 0.226.1
GOLDEN_SHA256 70ed8e9377dec22a9b493e55f222b0e25a49d7f3caec8c506e0412fd6baefe69
size 657 197 592 B
baked gitea.dooplex.hu/admin/felhom-controller:0.226.1
MinAgent 0.129.0 (read from the controller CHANGELOG header, not assumed)
THE EVIDENCE IS THE ROUND TRIP, NOT THE BUILD LOG. The published bytes were
downloaded back -- size and sha256 both identical to what the bake reported --
and ./etc/felhom-controller-image was read OUT of the downloaded archive:
`felhom-controller:0.226.1`. That is the delivered artifact naming the
controller it will start, from the bytes a customer's box would actually fetch.
Acceptance markers counted, not eyeballed, each string captured from this run's
own log rather than paraphrased from the runbook (two of the three the runbook
named until R-233 could not match anything the script prints): docker OK
(overlay2 = 1, including mount point rootfs = 1, mp0 = 1, upload OK (HTTP 201) =
1; excluding = 0, FATAL = 0, mp1 = 0. The 404 pre-gate passed before the run, so
nothing was overwritten.
THE VOUCH IS A THREE-FIELD CHANGE AND ALL THREE WERE CHECKED: agent_version
0.130.0 >= min_agent 0.129.0, so NOT the R-216 shape; wrapper_sha256 carried
through explicitly because the handler clears it when omitted. Verified by
RE-READING the manifest rather than trusting the flash -- golden option 0.226.1
SELECTED, all four shas matching.
THE FLOOR IS PROVEN ACTING, NOT MERELY SET. demo-felhom self-updated within 30
seconds: "[selfupdate] Post-update startup: update successful (0.225.0 ->
0.226.1)". Both demo machines now run 0.226.1 and only one of them was deployed
to by hand.
Token hygiene: copied file->file, read inside the VM by a runner script, never
on a command line (systemctl show ... | grep -c -F token = 0). THE LEAK GREP ON
THE COMMITTED LOG WAS PROVEN TO WORK BEFORE ITS 0 WAS BELIEVED -- a throwaway
copy with the token appended grepped 1, was shredded, and only then was the real
log's 0 taken as evidence.
Teardown: build guest 9100 destroyed --purge, secrets shredded AFTER the log was
copied out (standing rule 5), VM powered off, disk reverted to virgin.
R-242's OTHER half is untouched and still open: nothing gates the VOUCH itself.
One handler, two fields, opposite discipline. An unknown event_type is rejected
with a loud 400. An unknown severity was rewritten to "info" without a word -
and severityNotifies drops "info" before BOTH legs, so the event was stored,
answered 200, and mailed to nobody.
Two shipped features went out that way: DiskAlertKind.Severity emitted "warn"
until controller v0.215.0, app_start_failed until v0.223.0. Measured on the live
hub DB today: 91 app_start_failed events stored all-time, ZERO notification_log
rows before this session - not one, on any channel.
The mechanism built to catch this class was structurally blind to it: the
dispatcher's `unrecognized severity` line cannot execute for anything arriving
over the API, because the coercion one line earlier guarantees the value it
looks for cannot arrive.
The coercion STAYS - a rejected event is a lost event, and losing an alarm is
worse than mis-routing one. Only the silence is fixed: a WARN naming the
customer, the event type and the rejected value.
The dispatcher branch is KEPT, not deleted as dead, and the reason is evidence
rather than caution: cmd/hub/main.go wires dispatcher.ProcessEvent DIRECTLY as
the monitor.EventNotifyFunc for the staleness, host-staleness and offsite-box
checkers, which never pass through the handler. For those it is the only
severity guard there is. All 90 severity literals in internal/monitor are
already valid, so the guard is silent because the producers are correct.
Test count 702 -> 709. Red-proof seen failing: delete the WARN line and the
coercion test fails with "the hub rewrote a severity and said nothing".
Golden 0.223.0 baked and published (sha 9eaf39ac3921...), round-trip HTTP 206.
Vouching is the operator's act and was not done here.
The gate failed only on `released > baked`, so it could catch a forgotten bake
and nothing else. A golden AHEAD of the record passed silently - and that is
how controller 0.221.1 was built, baked AND vouched while the newest CHANGELOG
heading still read v0.221.0, with every gate green. Reproduced on the real
history: newest released 0.221.0 / newest golden baked 0.221.1 -> exit 0.
The gate now asks whether the version being shipped is WRITTEN DOWN: the baked
version must have its own `## vX.Y.Z` heading anywhere in the CHANGELOG.
Membership rather than `baked > released` deliberately - a comparison against
the newest heading alone goes green the moment any later entry is written,
leaving the unrecorded version permanently unrecorded. INCONCLUSIVE (exit 2)
preserved; every refusal names a reason and a route.
Red-proofed both directions: old gate/old record exit 0, new gate/old record
exit 1, new gate/fixed record exit 0, absent clone exit 2, post-bake exit 0.
08-alarm-ladder.md is new, and its absence was itself the finding: no document
owned "when does a broken app raise an alarm?". The rules lived as comments in
four packages, each locally correct, with the ordering between them legible only
by reading one function top to bottom - which is how R-384 survived review.
R-383 and R-384 closed into CLOSED-ITEMS with their rules kept. R-385 filed
closed. R-386 filed OPEN: a single-container app stopped out of band raises no
alarm, and a comment claims the opposite - measured live, 9 scans, 0 events,
against a positive control from the same box 17 minutes earlier. Not fixed here.
Golden 0.222.0 baked and published; vouching is the operator's act.
07-backup-architecture.md gains a dated [FACT] on R-361 - a comment asserting an
invariant the code did not have, for four months - and a [DESIGN] on the db_dumps
decision INCLUDING the trap it created: a stable list lets the already-current
early return fire, so per-capture housekeeping must sit above it.
00-capability-map.md records the NEGATIVE from Part 3 so it is not re-derived: a
held app does NOT raise the dead-app alarm. It aggregates to unhealthy, which
IsDownState excludes. Measured on the shipped build with the scans demonstrably
running over it. No suppression was built and no row opened.
R-383: the double-failure message names an undo copy that is not there - R-361's
own class, one surface over, observed on both 0.220.2 and 0.221.1.
R-384: an app whose database has died reads unhealthy and raises no alarm.
R-361 closed and compressed. OPEN-ITEMS 325236 -> 327266 bytes.
Golden 0.221.1 baked, published and round-trip verified. The golden-currency gate
blocked this push and that block is not circular, so it was satisfied rather than
bypassed - no --no-verify anywhere in this session.
07-backup-architecture.md 6.3 gains a dated [DESIGN] paragraph on replay ->
rollback -> hold, including why no engine flag closes it: --single-transaction
makes Postgres atomic, MariaDB DDL is not transactional, so the rollback is the
fix and the flag is a belt.
Drill record for the live walk, including the TWO defects the walk found in the
fix itself (a rollback into a re-created container; an operator route that
cleared the file while the running controller kept refusing) and the ONE
red-proof that PASSED, which is reported rather than omitted.
R-379..R-382 compressed into CLOSED-ITEMS.md. OPEN-ITEMS 330683 -> 325236 bytes.
STATUS.md restates the outcome and names the next operator step.
Baked in the drill VM per RUNBOOK-manual-build.md 4.0/4.1, carrying controller
v0.219.0 (R-356).
GOLDEN_VERSION 0.219.0
GOLDEN_SHA256 67b46f78f8ed9c7b1876265ab1bde9ec6798897898b1836acece9f3864a2aeb6
656832571 bytes
All five pass markers matched, both negative controls at 0. Verified by ROUND
TRIP - the published object downloaded again and its sha recomputed - not by the
number the script printed.
Both token-leak greps were proved able to convict before their zeros were
believed: planted copy grepped 1, shredded, then the 0 accepted.
Teardown complete: guest 9100 purged, four secret/script files shredded after the
log was copied out, qemu exited, disk reverted to virgin. The revert first
refused while qemu held the image, which is the runbook's own no-holder proof.
NOT vouched - that is a three-field operator save (golden_version 0.219.0,
agent_version 0.130.0, min_agent 0.129.0).
Both of the drill's HIGH findings are fixed in controller v0.218.0 and confirmed on demo-hp with
a negative control first — the same planted, hash-recorded fixture run through the same steps on
both builds.
R-355: paperless-ngx's PostgreSQL was dumped into a directory for a stack that does not exist, so
it never entered the recovery unit, the off-site copy or the restore; and because the same wrong
name reached writeSafetyDump, a destructive restore took no undo copy and the fail-closed refusal
was never reached. Fixed by reading the compose project label. Sweep proven able to convict
before its count was trusted: one affected app of 53.
R-354: the off-site restore had no named-volume leg. Now it replays them from the scratch unit,
before the database and inside the stopped window, and VolumesReplayed reaches the sentence.
The half-false comment beside the skip is corrected and the half that still holds is named.
Golden 0.218.0 baked and published, sha 8e427869d13eafb71562b77d1535eef6c7f32b4db24f659b988ec6d7db8f478b,
verified by round trip on the downloaded bytes. NOT vouched and the floor NOT raised — both are
the operator's decision, and raising the floor is what puts this on demo-felhom, which is still
on 0.217.0 and still has both defects.
R-367 filed: the dumps already written under the wrong name are stranded. Nothing deletes them
(an existing guard), they are adoptable by hand, and doing it automatically would be a migration.
Ceiling R-366 -> R-367.
GOLDEN_VERSION = 0.217.0
GOLDEN_SHA256 = 0276c5f638d140a861daba4ef25129896e259937eec0ae5cba7af42391315ad0
archive volid = local:backup/vzdump-lxc-9100-2026_08_21-21_40_05.tar.zst
MinAgent = 0.129.0 (unchanged from 0.216.0)
Acceptance markers counted in this run's own bake.log, not paraphrased:
docker OK (overlay2 1
including mount point 2 (rootfs and mp0 — there is no mp1)
upload OK (HTTP 201) 1
excluding 0
FATAL 0
Published package fetched back over HTTPS: HTTP 200.
Token handling: copied file->file, read by a runner script inside the VM, never on a command
line. systemctl show of the live unit contained it 0 times. The committed bake.log greps 0 for
the literal token AND the grep was first PROVEN to work on that same file by appending the token
to a throwaway copy (grep = 1) then shredding it — a 0 from an untested grep is not evidence.
Evidence copied off the VM BEFORE teardown. Then destroy 9100 --purge, shred token+runner+script
+log inside the VM (0 left), poweroff, waited for qemu using `ps -eo comm` (never `pgrep -f`,
which self-matches), and reverted the drill VM to `virgin`.
This unblocks the golden-currency gate, which correctly refused the previous push of the register
rows: "controller v0.217.0 is released and NO golden carries it". No --no-verify was used.
NOT DONE: the vouch. It is operator-gated and is a THREE-field change; vouching golden_version
alone would ship this controller onto an agent older than it declares it needs.
Closes the two-release day-0 gap that has been convicting CI since
2026-08-14. Run against RUNBOOK-manual-build.md 4.0 + 4.1.
GOLDEN_VERSION = 0.216.0
GOLDEN_SHA256 = ac004dc90d8cefccc5448377892f9cff3a4c3e1e27d0e11129120e38ac31c34b
archive = 656,970,239 bytes, controller image 0.216.0
template = debian-13-standard_13.6-1_amd64.tar.zst (listed live, not reused)
Baselines re-read on the machine and all four matched the sheet: controller
v0.216.0, its MinAgent 0.129.0, agent v0.129.0, previous golden 0.214.0. The
published agent artifact for the vouched agent_version was confirmed present
in the package registry rather than inferred from a CHANGELOG, and the R-216
check passed on the machine: MinAgent is EQUAL to, not above, the newest
published agent.
Verified beyond the script's own claim: the artifact was downloaded back out
of Gitea and hashed, and it matches GOLDEN_SHA256 exactly. A script printing
a digest and the registry serving those bytes are two different claims.
Pass markers (corrected post-R-233 list) all present, quoted with line
numbers in pass-markers.txt; excluding/FATAL absent; there is no mp1.
Token never reached a command line: copied file->file, read inside the VM by
the runner. systemctl show grep = 0. Token-leak grep on the COMMITTED log run
with its positive control FIRST -- seeded copy 1, real log 0 -- because a
grep -c that matches nothing also returns 0.
Teardown: guest destroyed and purged, token/runner/script/log shredded AFTER
the log was copied out, qemu exit confirmed with ps -eo comm (not pgrep -f),
disk reverted to virgin.
Vouched by the operator; verified by reading the hub's own store: golden
0.216.0 / agent 0.129.0 / min_agent 0.129.0, and the hub's recorded sha256
matches the independently downloaded artifact. That check was necessary
because golden_currency_gate.py says of itself that it checks the BAKE, not
the vouch.
repo_gates.py --fast now rc=0, all nine gates OK -- first fully green run
since 2026-08-14.
Capability map deliberately NOT changed: the day-0 row cites drill documents,
and the map's only golden literal is a dated historical citation on the
recovery-journey row which bumping would falsify.
R-334 is closed in a follow-up commit quoting this push's CI run id, since
closing it without one would leave the ambiguity a third time.
ListSupersededEscrow had zero production callers for nineteen days. It is the only
reader of a retained identity_blob, so the retention shipped in v0.93.0 was
material the product could not reach - proven on the fixture 2026-08-12, where a
code that opens a retained package was answered as a code that opened nothing.
New GET /api/v1/hosts/<id>/escrow/retained: self-scoped exactly as the current-row
GET, same recovery-mode gate, same audit event written BEFORE the bytes leave,
capped at 16. Rows with a NULL identity_blob are WITHHELD and returned as
unopenable_count - they retain the PBS key, not the repository password, so they
can never open what the caller is asking about, and serving them would let the
screen claim an earlier package is openable on exactly the boxes the original
defect hurt. The count is returned because their existence is load-bearing and
underivable by the caller.
The trade, stated rather than waved through: the hub still cannot read any of it -
sealed bytes in, sealed bytes out, no decrypt path, no recovery code ever held.
What widens is volume, bounded by self-scope, the recovery-mode gate and the cap.
The response is a NAMED TYPE, not a map, so the wire-contract gate can resolve it;
the wire is declared as a fourth ROOT and the gate now checks 182 tags rather than
174. A positive control shows that check is name-presence, not decodability -
filed as R-315 rather than reported as coverage.
Also: golden 0.214.0 baked, published and round-trip verified; the countdown on
demo-felhom cancelled on the operator's ruling (R-307); the spike that halted
Part 3 recorded as R-312; the set-aside store found unrecoverable as R-313.
Six hub tests through the real endpoint; four red-proofs asserted applied.
Closes R-296 (verified: shipped in v0.212.0) and R-301 (premise confirmed, fixed
in v0.213.0). Files R-302 with WHY the obvious condition was rejected, and R-303
for the missing markOrphaned guard - the co-render is now harmless, not
impossible. Bake evidence for golden 0.213.0.
GOLDEN_SHA256=8593516889eb93fe1691410d7306be8cb87ee835b8d2378740eb34022272f849
Round-trip verified on the served bytes. Not vouched - that is the operator's.
Four defects of one family, all shipped today: something the box already knows, thrown away or drawn
as its opposite. Agent v0.128.0, controller v0.210.0. NO HUB CODE, no hub bump, no ArgoCD sync.
R-265 (this repo). timeout-minutes: 5 on the gates job — every honest run in the observed session
finished in 18-34s, so this is ~9x the slowest and far under whatever reaped run 264 at 834s with no
log. The alarm mail now carries Elapsed (start stamp via $GITHUB_ENV; an absent stamp prints
"unknown (no start stamp)", never a bogus 1.7-billion-second figure) and its "names itself in the run
log" sentence is qualified so it cannot mislead when there is no log.
⚠ THE UNKNOWN IS NOT CLOSED. Whether the if: failure() alarm fires for a REAPED job is still
unverified. The timeout makes the reap unreachable in practice; it does not answer what happens in
one. Demonstrating it means deliberately hanging a run on main, which would leave the branch red for
a parallel session. Said in the workflow comment, the changelog, R-265 and the report — none of them
claiming it is answered.
GOLDEN 0.210.0 baked, published, round-trip verified, NOT VOUCHED. The currency gate went red the
moment the controller was bumped — correct — and is closed by the bake, never --no-verify. No
--no-verify anywhere this session.
⚠ THE AGENT WAS NOT PUBLISHED UNTIL THIS SESSION CHECKED, AND IT MATTERED. R-221's fix is in the
AGENT, and a fresh install takes its agent from the Day-0 manifest. The binary had been hand-deployed
to felhom-pve and never published, so agent_version 0.128.0 was not selectable and a fresh install
would have received 0.127.0 — the golden would have carried the controller fixes and NOT the one the
headline defect needed. Caught by checking each Day-0 value was FETCHABLE rather than assuming.
Published from the live-deployed bytes, sha-verified across the hop first.
Registers. R-221, R-259, R-258, R-265 CLOSED. R-266 MINTED (READY): the failed root statfs still
travels to the hub as a 0-of-0 disk; ranked LOW because it is the quiet direction — it can only miss
a true alarm, never raise a false one — and it is now a two-repo wire change governed by G-1's gate.
Highest ID moved R-265 -> R-266.
CONTEXT S-39 rules the convention this project was missing: "we do not know" is never drawn as
"fine", and the codebase has ONE way of saying it — an explicit ...Known bool companion checked in
the template. ROADMAP G-3 was explicitly blocked on that decision and is unblocked; what remains
there is a survey-and-convert of existing sites, not the gate.
Capability map row 93 CHECKED and it was NOT claiming something untrue — it is about the operator
notification path. But its narrative ("the page you open to ask whether ONE app is backed up")
invites the wrong reading, and the adjacent thing WAS false until v0.210.0, so the row now records
that the two halves disagreed and only the operator half was true.
Six red-proofs across the two code repos, each with the mutation asserted applied. The one that
matters: Part 1 Scenario A FAILED against today's tree, with the intended message.
Part 1's operator-present live validation is OWED and is the session's STOP.
repo_gates --fast: all 8 OK.
Recorded on arrival as 0.207.0, re-read from live hub_settings at the end and it is 0.208.0 — the
operator acted while the session ran. The ask is therefore 0.208.0 -> 0.209.0, not 0.207.0 ->
0.209.0, and STATUS.md plus the golden evidence now say so.
Caught only because the state was re-read rather than carried forward from the arrival note. A fact
recorded at the start of a long session is a fact about the start of the session.
CI convicted ALL 174 checked tags while the pre-push hook was green. Cause, read from the run log
rather than guessed at the second attempt: the search used `grep -rnE --include=…`, and the CI
runner's image carries python3 and git and deliberately little else — its grep does not support
`--include`, so stdout was empty and the gate read empty as "the tag is absent".
That is a gate silently treating a tool failure as a finding, which is worse than no gate, and it is
exactly the error-swallowing this repo forbids. A green from it would have been just as untrustworthy
as the red.
Fixed by removing the dependency, not by working around it: the search is now pure Python — one
token index per receiving repo, built in a single pass, no subprocess. Faster too (one walk instead
of ~350 greps), and unreadable-file / empty-repo cases now exit 2 INCONCLUSIVE rather than reporting
absence.
THE BEFORE CAPTURE WAS RE-VERIFIED, NOT RE-GENERATED — the stronger claim. All 40 fields recorded in
BEFORE.md were re-tested against the new implementation: agree=40, disagree=0, i.e. exactly the four
this session fixed are now present and the other 36 still absent. The number 40 stands under both
implementations; only the mechanism changed. The whole-token property survives by construction — a
token index treats `healed_at` and `privsep_healed_at` as distinct tokens.
This is the THIRD instrument defect this gate's own controls caught before it was trusted, after the
substring false negative and the dr_recipe over-opacity. The first two were caught by re-finding the
known instances; this one by the CI-versus-hook disagreement the workflow's alarm mail explicitly
says outranks whatever the push was for.