The volunteer guide has an English twin. It is a TRANSLATION, not a rewrite: 16
sections in the same order, identical step counts, table rows and warning blocks per
section (measured, 0 sections differing in structure). Word counts are NOT a twin —
English runs 19 % longer overall and up to 42 % on the short sections, because
Hungarian is agglutinative; the +-15 % criterion the task asked for does not survive
contact with this language pair, so structure is the measure reported instead.
Golden 0.258.0 baked, published and vouched, with its record. One run, no aborted
attempts: the 0.246.0 bake's two traps were both avoided by following its own record.
Token proven not to leak with a planted control before the zero was believed.
The waiver is NOT retired, and the record says why in one line: it is the mechanism of
operator ruling R-468, not a note about this golden, and deleting it would turn the
next release without a bake red immediately. It is also not load-bearing today.
R-595: the catalog's copy gate could not run in CI at all — six pushes red, six alarm
mails, while the local hook was green. Found by reading the operator's inbox, not by
anything in the session that caused it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Part C shipped the same day the pilot was read: fifty apps in three pushes, 1 031 of
1 032 strings. The English Apps list shows ZERO Hungarian app descriptions across all
53 apps — the only Hungarian left on it is the "Naprakész" badge (R-589) and the
language picker naming itself, which is correct.
The Hungarian Apps list is identical to the pre-slice capture once the per-session
CSRF token AND Docker's own "Up N hours" container string are normalised. Both
normalisations are stated in the evidence rather than applied quietly — the second
one moved because two hours of wall clock passed between captures, not because any
copy changed.
Fleet floor raised to 0.257.0 with the declared MinAgent 0.131.0, above the vouched
golden so the declaration carries it. demo-felhom went 0.255.0 -> 0.257.0 by itself
in under 12 seconds and THEN rendered the English tagline: the floor delivered the
feature, not a version string.
Rows: R-593 (papra describes a session-signing key as "the app's subdomain" — the one
string left untranslated) and R-594 (the catalog gate can convict a retrieval promise
but has no way to REGISTER a true one, which the shared vocabulary's design calls
for). R-560 closed. 281 -> 287 rows.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
10-localisation.md §7 goes from [DESIGN, proposed] to [FACT], with the numbers it was
specced against corrected — and one correction chose an instrument rather than a
footnote. The catalog has 1 032 copy strings, not 835. 832 carry a Hungarian letter,
which was right. But the ASCII-ONLY Hungarian is ~120 strings, not three: „Aldomain"
appears 53 times and „A szerver domain neve" 53 times, and the three the plan named
(„Igen"/„Nem"/„Nincs") do not occur in this catalog at all. An accent-only gate passes
every one of them inside an English block — R-565's blind spot arriving again in a
different repo.
New §10.6 records what was proven live rather than reasoned about: the English pages
show English; the seven Hungarian pages are byte-identical before and after the push
apart from the per-session CSRF token; and a 0.255.0 box with the block synced onto it
renders identically and logs no warning, in a 93-line window that contains the sync's
own lines, so the absence is evidence and not a dead log.
Rows R-589 (the update badge is Hungarian on an English page), R-590 (the data-folder
card's backup promise, likewise, and it is a promise about the customer's files),
R-591 (Stack.Copy() deep-copies five Meta fields and not the new I18n map — safe
today, which is precisely why it is a row), R-592 (three defects inside the new
catalog gate, closed the same session, each found by its own decoy). R-560 updated:
Parts A and B done, Part C waiting on the operator's read of the pilot.
STATUS asks for that read, and for the floor to 0.257.0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Live at iso.felhom.eu, sha256 dceacae5da247d76cad065bf6c0d3bbefac8d8a5f8e571
db2fd8449a97e94829, and both download pages now name it.
Every gate criterion is recorded with its OBSERVED value in
documentation/tests/iso-release-1.29.0-2026-09-18/ — including two proof
installs from the published bytes, one per boot-menu entry, each with a first
boot AND one reboot: /etc/issue bilingual with zero hits for 8006, pvebanner
masked, package 1.29.0 installed, unit enabled and fired, pairing code present,
and the installed script byte-identical to repo HEAD. G11: the downloaded bytes
hash to the published checksum.
A defect was caught BETWEEN builds by looking at the screen rather than at the
config: the second menu entry read "Felhom telepítés (szöveges mód) / Install
Felhom (text mode)" — 58 characters — and the GRUB menu box cut it at "Instal".
The English half was unreadable on the boot screen. Shortened to "… / text" and
rebuilt; the published image is the rebuilt one. The Hungarian half is the part
that may not change, so the English half is the part that gave.
Teardown: VMs 323/324/325 destroyed, the two unclaimed appliance registrations
discarded (zero left in `registered`), guest 9201 untouched — 23 containers
before and after. The two stale *.rootpw.txt files were shredded from the
publish source directory before the upload ran from it (R-587, files gone; the
guard that would stop it recurring is still open).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The image is BUILT but NOT PUBLISHED, and that is deliberate: publishing to
iso.felhom.eu is public and irreversible, and the runbook needs a proof install
on BOTH menu entries plus a reboot against the uploaded bytes. That is
supervised, so I stopped there. felhom-installer-1.29.0-pve9.2-1.iso, sha256
c67ceaa3…fb02, with every mechanically checkable criterion passing (G1, G2, G5,
G6, G7, G9, G16 — including both payload files byte-identical to repo HEAD).
The download pages still name 1.28.0, the image that IS published. Pointing
them at a file that is not there would hand every reader a 404. A new site gate
refuses the two pages naming different installer files or checksums, so
whoever publishes 1.29.0 cannot update one and forget the other.
Measured rather than read: the pairing banner is 24 rows on a 25-row console.
One row of margin — so the height is now pinned, because two more lines push
the HUNGARIAN code at row 5 off the top, and a banner whose code has scrolled
away is furniture.
R-587: two root-password files from July sit in the directory the public ISO is
published from. Both 404 on the bucket (against a 200 control), so nothing
leaked — but the only thing keeping them off is an --include pattern they miss
by an accident of naming. A pattern that protects by coincidence is not a
control.
R-588: release records live in two different places, which made me wrongly
conclude 1.28.0's gate had never been run. It had.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
THE IMAGE IS NOT BUILT AND NOT PUBLISHED BY THIS COMMIT. Publishing to
iso.felhom.eu is public and irreversible and its runbook requires a proof
install on BOTH menu entries plus the 16-criterion gate run against the exact
uploaded bytes. That is the operator's step. The download pages therefore still
name 1.28.0 - the image that is actually published - and a new site gate
refuses the two pages naming different files or hashes.
Three texts a person meets before any dashboard become bilingual: Hungarian
block first, byte for byte as before, then English, inside the same frame. The
pairing banner, the bound banner, /etc/issue (and the postinst's byte-coupled
copy), plus an English half on the GRUB entries.
The Hungarian is a GOLDEN, not a grep: test/golden/*.hu.txt were captured from
the script at 183727db9c before one English line existed, and the harness
asserts each banner's first N lines are exactly the golden. Red-proofed by one
changed byte, by an "a" planted in the English block, and by an over-wide line.
R-586, found on the way in: running the harness UNCHANGED at the base commit
failed two R-496 checks. The script paints with `>`, which truncates a FILE but
is a no-op on a console device; ISO 1.28.0's new bound banner (c033b3b) paints
straight after the pairing one and wiped it before the check read it. c033b3b
did not touch the harness, and nobody saw it because the harness is in no gate
and no CI run. Fixed with a FIFO; production code untouched. The harness being
ungated is still open.
The release gate's G16 required every Felhom string to be Hungarian and would
have STOPPED this publication. Operator ruling 1b of 2026-09-17 supersedes that
scope, so G16 is rewritten rather than waived: Hungarian FIRST, pinned by the
golden, each secret named once per language.
letoltes.html changes by four lines only. The English link is not in the nav -
the nav is a shared block site_gates.py pins across every page, and the gate
convicted the first attempt.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Part B's live proof is a LOG LINE rather than an e-mail, and the evidence page
says why: severityNotifies drops `info`, so every safe converted producer mails
nobody, and every converted producer above `info` describes something bad that
is not true. Triggering one would mean a false record on a real box's timeline
or a state change the fences forbid. The log line was built in v0.256.1 for
exactly this, after finding there was nothing to look at on either side of the
wire.
15:07:28 controller_started [hu-only] — Controller elindult (0.256.1)
15:08:25 controller_started [+household(en)] — Controller elindult (0.256.1)
The Hungarian sentence is identical in both, and the hub stored the Hungarian
in every case including the English-household push.
R-558 CLOSED, R-555 closed with it. R-585 filed: six producers still send
Hungarian only because their sentence arrives already finished from another
package — `offbox_enlarge_blocked` matters most, since it has no hub entry so
its raw sentence IS the household's whole mail.
10-localisation.md gains 10.4; STATUS rewritten for the operator with the two
decisions left (raise the floor to 0.256.1; whether to rotate the demo
password after R-584).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The live proof: the box's own "send test notification" button pressed twice,
74 seconds apart, on demo-hp. Reporting `en` it produced "[Felhom] Test
notification / Dear Customer, ..."; switched to `hu` it produced "[Felhom]
Teszt értesítés / Kedves Ügyfél! ...", byte-for-byte the v0.117.0 literal. The
operator's copy is identical in both, which is the half worth stating.
R-583 (closed, hub v0.118.1): the test mail was the one customer mail that did
not follow the language, and it is the mail an operator would use to CHECK
that the language works. The surface you would use to check a feature is the
one most worth checking first.
R-584 (open, P2): five probe scripts from slice 2's releases B/C/D were still
in the guest's /tmp carrying the controller password INLINE. The rule to
delete them exists, was loaded, and was not followed three times running - so
the rule is not the mechanism. All shredded; whether to rotate the shared demo
password is the operator's call.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found during v0.118.0's own live proof, which is the only reason it was found:
I went to press "send test notification" for the English demo box and read
sendTestEmail first.
It had its own hardcoded Hungarian subject and body and never went through
FormatCustomerEmail, so it was the one customer mail v0.118.0 did not localise
— and it is the only customer mail an operator can trigger on demand, which
makes it the one most likely to be used to check whether the localisation
works. Pressing the button for an English household would have answered that
question wrongly, and convincingly.
The two sentences are extracted byte-for-byte into the bundle, so the Hungarian
test mail is unchanged. Red-proofed against the hardcoded version.
The general form worth keeping: the surface you would use to CHECK a feature is
the one most worth checking first. A broken instrument that reports success is
worse than a broken feature.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The hub has written every customer e-mail in Hungarian whatever the box was set
to. The box has published its language since controller v0.247.0; nothing read
it. Now it does.
Nothing an operator reads changes. The Hungarian mails are byte-identical, and
that is a diff rather than a reading: 56 goldens per language captured from
v0.117.0 BEFORE any string moved, and all 56 Hungarian ones pass unchanged after
every sentence was routed through the new bundle.
- internal/i18n: flat bundle, 79 keys, hu authoritative + hu fallback, ceiling 0.
- customerMessages/severityLabels are DERIVED from the bundle, so a sentence is
written in one place and all 40+ tests that read those maps still work.
- Language order: last reported -> created-with -> hu. reports.language defaults
to EMPTY, never hu: "never told us" is not "chose Hungarian".
- message_customer on POST /api/v1/event, additive and optional forever, for the
sentences the box composes and the hub cannot translate.
- The bind page is per-language, and its `expired` state stays Hungarian: it is
the state an unknown token lands in, so rendering a real English customer's
token in English would make the LANGUAGE answer what the TEXT refuses to.
Two defects found inside the release:
- R-581: the newest report was picked by received_at, which has SECOND
granularity, so same-second reports tied and the winner was arbitrary. Ordered
by the autoincrement id now. GetCustomers() still has the shape - row open.
- R-582: the English copy-guard stems, ported word for word from Hungarian,
convicted 141 honest sentences. The English claim is a phrase with a modal.
R-555 closed: the language allowlist entry is out of wire_contract_gate.py.
hub_copy_gate.py follows the sentences into the bundle - without that it would
have scanned four files that no longer hold any customer text and reported
success. Three new decoys incl. an innocent control.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
POST /configuration/global-floor with min_controller_version=0.255.0 and the
declared min_agent=0.131.0 (R-472); 303 flash=floor_set, read back from the
form, not from the POST.
The proof is the N100: it was never hand-deployed and its own Docker reports
felhom-controller:0.255.0 healthy within five minutes of the save. Three boxes
remain below — all BLOCKED or DOWN, which is a floor being held, not a floor
failing; each takes it on its next check-in.
R-580 filed: curl's %{redirect_url} rebuilds the request URL WITH the --netrc
credentials in it, so the hub password was printed into the session's own
output. Nothing written to a file, nothing committed. The build-deploy skill
now carries the rule.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-579 filed and closed the same day: five shell templates loaded style.css
with no cache-buster, so a browser holding the pre-0.254.0 file rendered the
new globe unstyled; and the globe sat outside the card.
- STATUS.md rewritten for the operator: what was seen, why, the third defect
found while fixing it (version disclosure on the guest share page, caught by
TestShareGuest_HeadersTilesNoAdminChrome), and the one decision left —
raise the fleet floor to 0.255.0, with what happens either way.
- 10-localisation.md §3: the shells' asset tag, and why the two guest pages get
an opaque tag rather than the version.
- Audit D: the parity diff (91 of 106 fixtures identical, every dashboard page
among them) and the live endpoint evidence from demo-hp guest 9201.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Agent versions checked FIRST (both live boxes 0.132.0, above the declared 0.131.0), because a
floor is held for a box whose agent is below the requirement. Positive observables at both ends,
and the box's is the one that counts: demo-felhom's own log reads "settle-gate: GO — at/above
floor 0.254.0", its image file and running container agree, and its four other containers stayed
up. The release's visible change is on that second box too: one globe on the sign-in page, zero
of the old text links.
What the table must not be read as: the two DOWN customers got nothing and will take 0.254.0
unattended when they next report, from 0.115.0 and 0.245.0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
10-localisation.md §10.3: the saved notes follow the box language at write time, with the
one-night consequence stated rather than hidden; the globe, and the table of WHO reads which
page and where its globe posts — getting that wrong makes the button do nothing, which it did
on /recovery until the live probe found it. Decision 6 superseded a second time; decision 8
(a claim carries the visitor's language) recorded. Decision 5 of §11's anonymous-surface line:
changing what a VISITOR reads is within what an anonymous request may do; changing anything the
household owns is not, and POST /lang can do only the first.
R-578 — the deadlock, and why it is a row rather than a fixed bug: UpdateOffboxStatus holds the
settings write lock while running its callback, boxLang() wants the read lock, sync.RWMutex is
not reentrant. On a real box an off-site run would have hung FOREVER holding that lock. The
symptom was a test suite going from 8 minutes to a 25-minute timeout. Fixed and guarded, but the
guard covers one package and three helper names; the class needs a gate.
R-577 — a guest share visitor still has no way to pick a language, and the household's setting
is the wrong default for a stranger. Deliberately left, pinned by a test, and the operator's to
decide because it is a promise the share feature makes.
.claude/rules/live-probes.md, unconditional: never send a deploy request for an app that is not
installed, not even expecting a refusal — the endpoint accepts first and validates later. Two
sessions made that mistake in two days, the second WITH a prompt line forbidding it. A prompt is
read once; a rule file is loaded every session.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The floor went 0.250.0 -> 0.253.0 with min_agent 0.131.0 declared, which is what carries a floor
above the vouched golden (R-472). Agent versions were checked FIRST, because the floor is held for
a box whose agent is below the requirement.
Positive observables at both ends, and the box's is the one that counts: the hub logs "managed
floor SERVED … from declared", and demo-felhom's own log reads "settle-gate: GO — at/above floor
0.253.0 (we are 0.253.0)" with the running image and four untouched containers to match. The
release then answered in both languages on that second box.
What the table must not be read as: the two DOWN customers got nothing and will take 0.253.0
unattended when they next report, from 0.115.0 and 0.245.0 — nobody has carried a box forward
from 0.115.0 in one step. Ordinary for a floor, and stated rather than left implied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
All 179 Hungarian error messages carry a key; zero remain. 10-localisation.md gains §10.2: the
four properties util.MsgError had to have at once and the failure each one prevents, and the
plural rule as a BUNDLE rule rather than a per-call-site flag, with the answerable sentence and
what the other option would have cost.
Two instrument defects recorded rather than tidied away, because both shapes recur:
R-576 — the parity gate has a measured blind spot. The bulk converter dropped the continuation
of multi-line concatenations, damaging 7 producers, and the gate stayed GREEN: every surviving
fragment WAS a byte-equal base-commit literal, so its question ("is this text real?") was
answered yes while the CALL had lost half its sentence. Two behaviour tests caught it. The
general form: a structural gate over the TEXT cannot see a defect in the CALL.
And the script counting what was left was case-sensitive, so it said "0 remain" while five did —
R-565's shape inside the measurement. Every "no Hungarian left" claim in this slice is now made
case-insensitively and with both controls.
R-575 — the soft memory-overcommit warning has no error to carry a key and no language where it
is built, so it renders Hungarian on an English page. Named in the code, not hidden.
Live evidence includes a mistake I made and corrected: a probe of the deploy refusal INSTALLED
vaultwarden on demo-hp (the endpoint accepts before it validates), the same mistake the previous
session recorded. Removed through the product's own path with its data; verified gone.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Endpoint-level on demo-hp guest 9201, method stated. The flash key renders the sentence in
each language where 0.251.0 showed the raw key; a link an OLD controller minted still shows
its prose, in both languages; country names follow the language and the list re-sorts.
Hungarian parity measured on the page BODIES, not on a hash: ten pages on 0.252.0, the guest
rolled back to 0.251.0, the same ten again, then rolled forward. Six byte-identical; the other
four differ only in live state (a clock crossing a minute, an app unhealthy for a moment, and
the update-available line, which is true on the old version and false on the new). The first
pass had saved only hashes — a hash cannot show WHAT moved — so the before-state was
reproduced rather than asserted.
Box left on 0.252.0 with its saved language `hu`; every English probe used the ?lang= override,
which is not persisted. Provisioned nothing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
10-localisation.md gains §10.1: the measured numbers (1 120 base literals, not 1 141; 226
converted), the flash-as-key design and why a key must be resolved by the READER, word order
through Go's explicit argument indexes rather than a second placeholder syntax, and the parity
gate that makes "byte-identical" a measurement instead of a reading.
Two claims the plan carried that live source disproved, both about the wire, both recorded:
the country table is NOT on the wire (only codes are), and the hub does NOT always compose its
own customer mail — it falls back to the controller's event message, which is why those 31
sentences stay Hungarian until R-558. That survey is handed to R-558 as its input list.
R-566 CLOSED (four app-named page titles now carry a %s). New rows R-572 (two funcmap helpers
with no English form), R-573 (the two channel-health banners arrive as finished Hungarian),
R-574 (handler_debug.go mixes page copy with payload).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- audits/r553-2026-09-17/: live before/after on demo-hp (CSRF redacted), the 409 refusal, the hub
health block, red-proofs for all five sites, the site-5 fixture diff, gates.
- 10-localisation.md §9: the five decisions with what each reads now; the rule (a text signature may
remain only where the text is not ours); the one legacy exception and its end date.
- Register: R-553 and R-563 closed to CLOSED-ITEMS; R-569 (four API handlers match English words),
R-570 (the legacy stale-note fallback + the slice-2 fence), R-571 (classifier and alert placement
documented nowhere). R-557 carries the R-570 dependency. 263 -> 264 open.
- STATUS, including the live probe that installed an app and was removed the same minute.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Design written after the spike ran (controller v0.247.0 live on demo-hp): mechanism,
flow, fallback, gates per language, catalog model, sliced plan with costs, operator
rulings 1-4 recorded, CC decisions 5-6, open decisions 1b and 7.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Phase 0 of the localisation starter: i18n_inventory.py counts every customer-visible
Hungarian string; the audit names six further claims in the prompt that live source
disproved. Rows for the compare-not-show sites, wizard deletion, the wire-contract comment
blind spot, and localisation slices 1-6.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Decision 1 delivered: floor 0.246.0 served, the N100 on 0.246.0 within seconds;
Peti's box is DOWN on the hub and receives it when it reports.
Decision 2: the agent vouch was refused by R-120 until a newer golden existed.
On the operator's choice, golden 0.246.0 was baked (sha 05b7559d, amd64, all
markers, token leak 0 with control 1, registry 200 before teardown) and vouched
together with agent 0.132.0. golden_currency_gate: WAIVED -> OK. Bake evidence
filed where the gate and runbook read it: tests/golden-0.246.0-2026-09-17/.
Recorded, none reaching the registry: a first attempt on the arm64 template
(my version sort), a self-matching pkill, and an OOM-killed watcher whose
post-bake steps were done by hand.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Decision 1: global controller floor 0.244.0 -> 0.246.0 with declared MinAgent
0.131.0. Served for demo-felhom; the N100 ran 0.246.0 within seconds; demo-hp
already did. Peti's box is DOWN on the hub and cannot receive it until it reports.
Decision 2: vouching agent 0.132.0 was refused by the hub's R-120 gate - the
vouched golden 0.245.0 is older than the newest controller the fleet reports
(0.246.0). Nothing stored. A golden >= 0.246.0 is needed first; put to the operator.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-539 closed: five real controller kills on demo-hp 9201 with the production
24 h window raised controller_slow_crashloop, and exactly one operator mail
arrived (09:29:40Z). The fast brake never armed.
Register 215 -> 213 open (R-551, R-552 filed; R-539, R-546, R-549, R-550 closed).
unproven.py: 35 of 55 not walked, no number moved.
The report names the brief's wrong claims first and one recommendation not
followed: the controller floor was not raised - validated on one guest, a gap
of my own found during validation, and a floor above the golden reaches Peti's
box too. The operator's call.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Architecture: 08 records ruling A (45 m, the round-9 arithmetic, the cost) and
the two new event types with their audiences; 03 records the slow counter as
built; 07 records the restore-record persistence as a REVERSED design for the
restore record only; CONTEXT.md carries the day's rulings.
Guide: the recovery code moves after the first apps and waits for the yellow bar.
Register: R-549 and R-550 closed PROVEN-LIVE; R-546 closed on red-proofed tests
with its live walk owed by R-551 (no Tier-0 box is paused AND agent-connected).
R-552 filed: an interrupted-restore notice for a removed app never clears -
found in my own v0.246.0 after the release was built.
Evidence: Part A (hub prints 45m/1h30m), Part C delivery on HP and N100, B.4(a)
live proof and its teardown.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Three report cycles instead of two. The chaos-night round-9 measurement: cadence
15m, a failed push retried for ~100 s, a 29m59s gap against a 30m threshold -
one second from paging the operator about a healthy, self-repaired box.
node_down and host_down move to 90m with it (2x). The same release makes the
dashboard's customer status read this value instead of a hardcoded 30m.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-549 (operator ruling A): controllerStatus hardcoded 30m/1h while both
staleness checkers and hostStatus read alerting.stale_threshold. Moving the
threshold to 45m would have painted a customer amber 15 minutes before the
alarm could fire - the second definition rollup.go's header forbids. It now
reads the same value, down at 2x. Both 'checker initialized' log lines print
the threshold, which no line did before.
R-539 (ruling 3 of 2026-09-16): controller_slow_crashloop (warning,
operator-only), minted when the agent's slow_crashloop_since moves, with the
fast sibling's first-sight rule.
R-550: restore_interrupted (warning, for the household) allowlisted with a
Hungarian customer message.
Red-proofs, each seen failing then passing: the status test with the old
hardcoded numbers; the checker test with the movement branch removed; the
operator-only test with the registration removed; the household-message test
with the Hungarian entry removed (asserted on the SUBJECT - the body
legitimately repeats the raw message, which my first version of the test
mistook for a fallback).
go build/vet/test ./... green, 18 packages.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Eight per-round accident logs (rounds 3, 4, 5, 7, 8, 9, 10, 11) sat untracked
in the evidence directory, plus round 1's poll log's final three lines - the
TCP reset that ended the ghost task. Found by the clean-tree check at the start
of the next task. Evidence, no content change to any finding.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
STATUS.md gets the morning note in the rules' order - decisions (none under the
unattended rule), what was exercised, what broke (nothing in the product; three
fixes worth making, all filed), rows (five opened, none closed), what could not
be tested, cleanup, and what needs the operator with the cost of doing nothing.
REPORT-chaos-night-2026-09-17.md follows template section 15 and names the
prompt's wrong claims first. It is a separate file because REPORT.md holds the
earlier session's write-up and this repo's rule forbids clobbering it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The acknowledged delete went through at 07:25:13Z (confirm_host_id +
delete_escrow=1 -> 303). Every line of the after-state written down before the
act matched: the host 404s; drill-r50, both demo hosts and the tester-1
customer still 200; the customer lists zero hosts; ep0 identical across three
readings - six snapshots, 16G, nothing removed.
The automatic connect mail arrived two seconds later and is quoted with its
token redacted. It is provably tonight's: the mailbox held no such mail newer
than 18:17:46Z when checked at 00:38Z.
Why the hub layer finished six hours late is mine: the retry guard refused to
post while the host page contained the word ONLINE, and that word sits in a
JavaScript string that is always on the page. The hub's structured status said
'down' from about 00:54Z. It gave up at 01:20Z and nothing ran again until
07:24Z. Eleventh instrument fault of the night.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
drill-r50 and the tester-1 CUSTOMER record both captured at 200 while the
delete is still pending, with the four host records listed and the customer
page showing exactly one host - tonight's box.
The expected after-state is written down BEFORE the act, so it cannot be
adjusted to fit what happens: three host records left, the other three still
200, the customer record still 200 with zero hosts, and ep0 untouched. A fence
is only proven by a comparison, which is why ep0 was listed twice too.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The two the brief flagged were both TRUE and both checked tonight rather than
assumed: the automatic self-bind mail really was waiting (18:17:46Z, zero
operator presses), and the WG hook really does re-issue by itself after an
acknowledged delete - pbsdr_auto_reissue at 20:19Z, the F-14 path measured live
for the first time. Neither pre-declared press was needed.
The ones that were wrong: the off-site app restore was impossible on a rebuild
box whose repository is orphaned by design; the fixture ended with three disks
rather than two, which is my deviation and not the brief's; round 7's drawn
'update' could not run because the catalog's own canary failed, so 'use' ran
instead and was logged; 'an internet cut tests hub unreachability' was false
here because the hub resolves to a LAN address, which is my error against my
own recorded warning; and the schedule table's clock column was nominal - the
night's twelve rounds finished at 00:17Z, about four hours earlier than the
table suggests, with the drawn order, apps and accidents never changed.
Plus one the brief did not make and the night could not answer: the
dropped-event path remains unmeasured, because no event coincided with any of
the three hub outages.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Phase 2: all twelve front doors 200 and all 26 containers healthy (traefik and
cloudflared carry no healthcheck, so a bare Up is correct). The off-site
restore could not be done and two independent instruments agree why - restic
itself and the product's own status surface both say the repository is orphaned
with zero readable snapshots, because this box is a rebuild whose restic
password was minted fresh. The product surfaced that honestly within seconds.
What DID leave the house: the whole-guest copy on ep0, two intact snapshots
including tonight's 21:59:54Z one. Stated as a limit: that is a listing, not a
verification, and a PBS verify writes state so it was not run.
Household loop: 204 probes, 7 flagged, only 2 real events - both single-sample
outages during the two abrupt stops. Three were my own classifier counting a
301 as a failure, corrected in the log's own words. Ten of twelve rounds left
no mark, which is the 2-minute sampling rate and not proof of nothing.
Catalog bump verified reverted. Mail delivery proof recorded: raised and
delivered are two different claims and only one had evidence before tonight.
Interventions: 1 of 4. Both pre-declared presses unused - both prompt claims
they insured against turned out true, and the F-14 path was measured live for
the first time. Phase 0's seeding repairs listed separately because that damage
was mine; harness acts excluded with the reason stated.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Headline, three lines. Interventions: 1 - round 6's local backup leg, whose
off-site leg then succeeded unaided; both pre-declared presses went unused.
Ready for a volunteer: still yes - nothing cost a byte of customer data, the
box healed itself every time with no human, and all 17 alarms were true, none
missing, every one delivered. The pair that hurt most: restore + hard reset,
not because the box suffered (26/26 containers back in 150 s) but because it is
the only pair where the household is left not knowing what happened.
Capability map: a new PROVEN-LIVE row for a random night of household actions
under accidents, carrying what it does NOT claim - per-app off-site restore
untested (orphaned repo by design), the dropped-event path still unmeasured
because no event coincided with any hub outage, twelve rounds is a sample not
coverage, and the household loop's 2-minute sampling means ten rounds left no
mark in it.
The unaided-recovery-journey row gets a second scope note rather than a change:
tonight did not walk it and could not have, so its PROVEN-LIVE still stands on
0.206.0 only - neither re-proven nor contradicted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The mailbox closes a gap the truth table could not: 17 alarms fired and all 17
were true, but that was the FEED. The mailbox shows they reached a person.
Round 11's full set arrived - storage_disconnected naming the drive, four
app_start_failed naming exactly the four apps whose data was on it, then
health_degraded. Round 6's whole_guest_backup_failed names the TIER. Round 1's
offbox_repo_orphaned was mailed within seconds of the run.
What is absent matters too: NO node_stale mail for tonight's box after round
9's outage, exactly as R-549 predicts - the gap was 29m59s against a 30-minute
threshold. One second the other way and this would be a page-out for a healthy,
self-repaired box.
Teardown layer 3 began with a DELIBERATE un-acknowledged delete, to see the
gate refuse: HTTP 409, 'Host is ONLINE - deletion is refused'. Gate one fired,
not the escrow gate - the box died inside the hub's 30-minute liveness window.
The record survived, verified by its own URL returning 200 rather than by
counting substrings on a list page (grep -c counts lines, not occurrences - my
'3 then 2' was my error, not a deletion).
The wait is the product's and not mine to shortcut: the acknowledged delete is
armed for 00:55Z behind a guard that will not post while the host reads ONLINE.
The fence says the hub is never changed, so a liveness gate is waited for.
Also recorded: the mailbox baseline proving no connect mail exists from tonight,
so the one quoted after the delete is provably new - the exact check the brief
said had been skipped before.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Machine layer: VM 336 destroyed with all three disks purged, gated on its NAME
rather than its number because 9201 and 9202 share the host. qm list now shows
no VMs; /mnt/hdd_1/images/336 is gone; both standing guests still run.
Host layer, before and after: nvme-scratch 6.78% -> 1.61% (~48.5 GB released),
/mnt/hdd_1/images 59G -> 9.4G with only the scratch guest's own 9202 directory
left, free space 827G -> 875G. local-lvm UNCHANGED at 44.75% - the fence that
said 'local-lvm never' held. Firewall back at baseline with 0 physdev rules, so
none of the three network accidents left a rule on a host carrying two standing
guests. Both harness units stopped and disabled before the box died; the disk
guard's log was 0 bytes - it never fired once.
9202: nothing to remove, shown rather than asserted - three infrastructure
containers, 55 catalog TEMPLATES none of which was touched since 21:00, no
app.yaml marked deployed, no offbox config. A false label in my own transcript
is corrected there: I printed '(nothing listed above = no app stacks)' directly
beneath 55 names.
ep0: read again immediately before the delete and identical to the baseline.
The single-host delete handler shows no ep0 cascade, but a grep returning
nothing is the weakest evidence there is, so the store gets a before and an
after rather than an inference.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The before-picture that cannot be retaken once the machine is gone: pvesm
status, the guest list, VM 336's full config and the contents of /mnt/hdd_1.
Two things it records that correct my own assumptions:
* the machine has THREE disks, not the two the brief specified. The third is
the 64G disk I added during Phase 0 to extend the thin pool after filling
it with twelve simultaneous deploys. My damage, my remedy, and a deviation
from the fixture the brief described - declared rather than quietly torn
down.
* the /mnt/hdd_1 claim is now earned: nvme-scratch is defined with
path /mnt/hdd_1, is_mountpoint yes, and the three raw files sit in
/mnt/hdd_1/images/336.
The harness is stopped and disabled, its logs copied off first (R-320):
household 204 lines, diskguard 0 bytes - the guard never fired all night.
And a correction one minute old: I announced that the earlier log copy was
twelve lines short and that re-copying rescued them. It was not short - both
copies are byte-identical. I compared a line count read at 00:19 against a copy
taken at 00:31. Nothing was lost; only the accuracy of the record was at risk.
unproven.py: 35 of 55 not walked - NO NUMBER MOVED, which is correct for a
validation night that shipped no product code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The defect I almost filed: I suspected the household is never told the off-site
repository is orphaned, because 'elarvult' appears zero times in 60 KB of HTML
and my search was sound (negative control 0, two positive controls finding real
Hungarian text). Wrong measurement. The page's script fetches
/backup/offbox/status and the page carries offbox-orphan-card, orphan-reveal,
orphan-confirm and a triangle-alert icon. The household IS told, in a card
rendered client-side. Nothing filed - caught BEFORE the row existed, unlike
R-550.
The worthless probe: my attempt to read a verify state out of the ep0 manifest
returned nothing, and so did its negative control. With a compressed blob,
'no match' and 'unreadable' are indistinguishable, and I had no positive
control. So the verify state is UNKNOWN, not absent, and the only claim that
stands is that the copies are present and well-formed.
The catalog: verified reverted rather than remembered - clean tree, level with
origin, original redis pin and catalog_since intact, newest commit 2026-09-15.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
restic with the box's own credentials: 'Fatal: wrong password or no key found',
exit status 1. The product's own surface: orphaned true, snapshots 0, status
error. The repository is orphaned because this box is a REBUILD - its restic
password was minted fresh, so the existing snapshots cannot be opened. The
product surfaced that honestly as a true alarm in round 1.
So Phase 2's 'restore one DB-backed app from off-site' has nothing to restore
from. That is a fact about the fixture, not a product failure.
Recorded alongside: ep0 holds two intact whole-guest snapshots for this box,
including tonight's 21:59:54Z copy - round 6's off-site leg, the one that ran
by itself after I killed the local leg. Its file index is four times the size
of the afternoon copy. So the data did leave the house.
Stated as a limit, not glossed: that is a LISTING, not a verification. A PBS
verify would prove restorability and writes state, so it was not run - ep0 is
read-only for evidence tonight.
Tenth instrument slip recorded: my first restic probe printed 'exit status 0'
beneath a fatal error, because the zero belonged to the head at the end of the
pipe. restic 0.14 has sftp.command, not sftp.args - established by asking
'restic options' rather than assuming.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The row first claimed a restore leaves no record anywhere, citing four status
endpoints that 404'd. All four were paths I guessed. The real route, read out
of the restore page's own JavaScript, is /api/backup/restore-status and it
exists.
The corrected finding is narrower and better: the endpoint answers with the Go
zero value (started_at 0001-01-01T00:00:00Z) and carries no 'last' field at
all, while the page's own script renders '<operation> sikertelen.' from
st.last.message. The restore record is in-memory only and does not survive the
machine stopping - exactly the case a hard reset creates.
The original wording is left visible in the audit with the correction beside
it; the register row is corrected in place because a register must be accurate.
The reusable lesson: I found the real routes by asking the controller for its
own rendered links. Guessing produced four confident 404s that I then reported
as a property of the product.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
192 probes from 21:06:30Z to 00:18:10Z. Seven lines flagged as failures, of
which only TWO are real events - and both are outage pairs caused by an
injected accident: wiki during round 2's power cut, cloud during round 10's
hard reset. Each lasted less than one probe interval and healed by itself.
Three of the seven were my own classifier counting a 301 redirect as a
dashboard failure. The log carries the correction in its own words at
21:11:25Z, and the wrong lines were left in place so the correction is visible.
Stated rather than glossed: ten of the twelve rounds left no mark in this log
at all, including the twenty minutes with the drive pulled. The loop samples
each name every two minutes, so that silence is a limit of the instrument, not
proof the household saw nothing.
Ninth instrument slip recorded: the first listing printed nothing and said
'binary file matches' while the COUNT had already printed, so the summary looked
complete while the detail was dropped. Not corruption - zero null bytes, one
line of padding spaces. Re-read with grep -a.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Steady at t+1s, every door 200 at both readings, household 2 lines 0 failures,
and no new alarm - the newest feed entry is still round 11's recovery. Nothing
was raised about a box nothing was done to.
Checked rather than assumed: inject.sh's default branch exits 2 on an unknown
accident, so the control rounds never reach it - the runner handles the
no-accident case itself. A broken injector produces exactly the same result as
a control round, and only the code path distinguishes them.
Alarm truth table, all twelve rounds: 17 alarms fired, 17 true, 0 missing.
Three design gaps filed (R-547, R-549, R-550).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
ep0 baseline (read-only, nothing removed, no prune): three namespaces, each
with ct/9201 holding 2 snapshots, six in total, 16G used of 98G. The teardown
repeats this listing so 'the backups stayed' is a comparison, not an assertion.
Phase 2 readiness: scratch guest 9202 is running with a healthy controller and
restic 0.14.0 inside the controller container, but has NO off-site target
configured - so the restore will need the box's own repository address and
password handed to it.
Recorded decision: I did NOT read those from the box while round 12 ran. Round
12 is the closing control round and its whole value is that nothing was done to
the box during it. The read costs nothing to defer; the round cannot be re-run.
Also recorded: the eighth instrument slip - perl locale warnings plus a head
that cut the output before the marker returned an empty block that looked like
'9202 has no containers'.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The richest alarm round of the night, and every alarm was true and correctly
paired: storage_disconnected naming the drive by the household's own label,
four app_start_failed naming exactly the four apps whose data lives on that
drive, health_degraded, then storage_reconnected and health_recovered.
The other eleven apps kept serving throughout. The box recovered unaided in
67 s after the drive was plugged back in, with the front door following at
128 s. The drive came back clean: 98 G, 2% used, mountpoint config unchanged.
The round's own snapshot showed health_degraded with no recovery, which would
have been the first missing alarm of the night. The recovery had fired seconds
after the snapshot. A re-read taken after the precondition found it. Not a
missing alarm - a premature reading, caught by the discipline the earlier
mistimed readings forced.
Truth table now: 17 alarms fired, 17 true, 0 missing, across eleven rounds.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS