- 09 §6.1 phase table (copying, undoing, undone), §6.1a SHIPPED with the two
live-only defects, §6.4 part 1 SHIPPED.
- Capability map: a failed update is undone by the box - PROVEN-LIVE.
- Live evidence on 9202: three apps undone by the product with seeds before
the backup, after it and seconds before the press read back; cut-off copy
held honestly; power cut during the undo resumed; manual press after undo.
- Register: R-637, R-639, R-641, R-642 closed; R-638, R-640 narrowed; R-643
ruled; R-646 opened. STATUS asks the floor question.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- audits/r553-2026-09-17/: live before/after on demo-hp (CSRF redacted), the 409 refusal, the hub
health block, red-proofs for all five sites, the site-5 fixture diff, gates.
- 10-localisation.md §9: the five decisions with what each reads now; the rule (a text signature may
remain only where the text is not ours); the one legacy exception and its end date.
- Register: R-553 and R-563 closed to CLOSED-ITEMS; R-569 (four API handlers match English words),
R-570 (the legacy stale-note fallback + the slice-2 fence), R-571 (classifier and alert placement
documented nowhere). R-557 carries the R-570 dependency. 263 -> 264 open.
- STATUS, including the live probe that installed an app and was removed the same minute.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-539 closed: five real controller kills on demo-hp 9201 with the production
24 h window raised controller_slow_crashloop, and exactly one operator mail
arrived (09:29:40Z). The fast brake never armed.
Register 215 -> 213 open (R-551, R-552 filed; R-539, R-546, R-549, R-550 closed).
unproven.py: 35 of 55 not walked, no number moved.
The report names the brief's wrong claims first and one recommendation not
followed: the controller floor was not raised - validated on one guest, a gap
of my own found during validation, and a floor above the golden reaches Peti's
box too. The operator's call.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Architecture: 08 records ruling A (45 m, the round-9 arithmetic, the cost) and
the two new event types with their audiences; 03 records the slow counter as
built; 07 records the restore-record persistence as a REVERSED design for the
restore record only; CONTEXT.md carries the day's rulings.
Guide: the recovery code moves after the first apps and waits for the yellow bar.
Register: R-549 and R-550 closed PROVEN-LIVE; R-546 closed on red-proofed tests
with its live walk owed by R-551 (no Tier-0 box is paused AND agent-connected).
R-552 filed: an interrupted-restore notice for a removed app never clears -
found in my own v0.246.0 after the release was built.
Evidence: Part A (hub prints 45m/1h30m), Part C delivery on HP and N100, B.4(a)
live proof and its teardown.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Hub v0.112.0 serves a floor above the golden with a declared MinAgent;
controller v0.239.0 reached both demo boxes by that floor in 14 s and 15 s
and updates on any backup tier. 09 §3 decisions 7 and 8, §6/§6.1; 07 §6
line; capability map row; STATUS items 15/16 done and the cadence line
corrected; CONTEXT; register: R-470/R-472/R-475 compressed to CLOSED-ITEMS
(full text at 2f5d3af), R-477..R-480 opened, R-474 reproduced a third
time. OPEN-ITEMS 431689 -> 432156 bytes, CLOSED-ITEMS 118051 -> 120598.
Evidence: documentation/audits/rulings-r472-r475-2026-09-13/.
Controller v0.237.0-v0.238.1: the Update button is a guarded job — refusals, backup-first when the
proven Tier-2 copy is stale, safety dump, pin, pull (pin back on failure), health, HOLD on failure.
Proven live on demo-hp: A, B, E, F, H and the restore walk (audits/slice4-2026-09-13/).
Correction to this morning's pages: between golden bakes the hub HOLDS a floor above the vouched
golden, so a release does not reach the fleet by floor (R-472, operator decision). Corrected in the
runbook, STATUS, CONTEXT, R-468 and the gate docstring.
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-449. Until today one app upgrade out of 53 had ever been measured, by hand, and
the whole update arc was designed against that single data point.
C3 first: the negative control, whose TO image exits immediately, came back
failed. That is what makes the greens mean anything, and it cost 556s because a
negative is only honest if it waits out the full settle window.
Seven edges, three apps. All five real catalog upgrades kept the customer's data.
The finding that changes an assumption the arc was carrying: whether an upgrade
can be UNDONE is a property of the individual APP, not of upgrades. Docmost
refuses - 'corrupted migrations: previously executed migration
20260213T085259-notifications is missing' - and privatebin does not. That
reproduces the Nextcloud result on a second app by a DIFFERENT mechanism, so the
struck word 'rollback' now rests on two measurements instead of one.
The finding nobody was looking for, R-459: our own bookstack template moves
MariaDB across a major and sets no MARIADB_* env at all, so the engine logs that
the datadir upgrade it requires is being skipped, and serves anyway. The cause is
assigned rather than guessed - the app half alone produces no upgrade line, both
edges that move the engine produce it - which is exactly what decomposing E3 into
E3a and E3b was for. It also explains why E3's abort looked like it worked: the
datadir was never converted. Whether that ever breaks is NOT established, and the
row says so.
Also opened: R-460 (bookstack's file half cannot be seeded headlessly), R-461
(target-selection.md names a venue that does not exist and fences a VM that is
gone), R-462 (the widening, costed with this run's real numbers - and the cost is
dominated by fixtures, which do not amortise).
Teardown all three layers, hub checked rather than asserted. local-lvm read 30.50
percent before and after. The capability map was deliberately NOT edited: this
measured apps, not the product.
09-update-architecture.md gains the fourth dated operator ruling (2026-09-06,
Option 1) and its section 5 is rewritten from a proposed shape into the shipped
one: the pin, the stored definition, the render table, the four writers, the
startup ordering, and the trap this slice set for slice 2 - the live compose file
is now the frozen one, so a badge comparing against it would answer Naprakesz on
exactly the apps that are behind.
02-controller-module-map.md said 'copy compose + .felhom.yml'. That stopped being
true today, so it is corrected, and the two sections describing the old seam now
carry a banner saying they describe v0.234.0 and below - kept because every box
under v0.235.0 still behaves that way and because they are the measured account
of why it changed.
R-447, R-441, R-438 and R-455 closed and compressed into CLOSED-ITEMS; R-458
opened for the .felhom.yml asymmetry, with what would settle it by measurement.
Live evidence: two real catalog pushes travelling the real 15-minute cycle, both
reverted, the tree byte-identical afterwards. The restart that used to take 18.3
seconds and pull a new image now takes 0.1 seconds and pulls nothing.
THE RECORD WAS TELLING A WORSE STORY THAN THE TRUTH FOR TWO MONTHS, and my own probe is why.
R-429 CORRECTED. Yesterday's spike searched for a directory called `.snapshots`. The vendor documents
the path as /.zfs/snapshot. The probe's CONTROLS were sound and its SUBJECT was wrong, so "not found"
was true and meant nothing. Re-probed at the documented path on both boxes, with controls:
- /.zfs lists (shares, snapshot) from inside the jail;
- a write into /.zfs/snapshot is REFUSED - dest open ...: Failure - while the identical write to
the account home SUCCEEDS and was cleaned up.
That is the append-only property PROVEN rather than cited, and it is the sentence the whole re-scope
rests on. Seven daily snapshots are confirmed in the panel. The mitigation works. What remains true,
and was always the actual finding: the row claiming it had no R-number, its "confirm tomorrow" went
36 days unanswered, and the DUE-CHECKS block built for that class was empty. THE FINDING WAS NEVER THE
SNAPSHOTS - IT WAS THAT NOBODY COULD TELL.
R-95 RE-SCOPED: the box can delete its LIVE repository but cannot write to the daily snapshots of it,
so a deletion costs at most one day plus a per-file recovery - not open-ended loss. The ranking is
Viktor's; it has been #1 since July on the old story.
R-432 FILED: a sub-account sees /.zfs/snapshot EMPTY while the same box holds seven snapshots, so
per-file recovery is operator-only today. One panel read settles whether a NAMED snapshot can still be
entered, which would make it product-reachable.
R-431 SHIPPED. Third signal in OffsiteChecker. On the hub deliberately: a detector on the box is one
the deletion can silence. Threshold REASONED, not invented - over 12 898 reports every decrease lands
on ZERO and predates stats_known, and in the stats_known window there are none, so observed churn gave
nothing to calibrate against. Retention cannot halve a total; a mass deletion goes to ~0. Hence: more
than half, and at least 5. Guarded by StatsKnown (R-331), the declared State (R-204) and run success
(R-100 - whose lesson lives in this very file).
ACCEPTANCE: 9 009 real report points replayed through the detector produced ZERO alarms.
Three red-proofs run. The escalation one only became real after the first version was found HOLLOW -
it re-swept the same report, so the baseline had already moved and the latch was never consulted.
07 row 10's status is NOT moved: the write-refusal is measured, but the recovery ROUTE has never been
walked, which is what PARTIAL means.
THE CLASS, now a row: an instrument that matches a LABEL rather than the fact it names. Five
instances - R-410, R-400, R-378, R-419, R-94 - and EVERY ONE was found by accident, by someone
looking at something else. The gates enforce every other rule in this project, including the rule
that findings must be written down rather than left in prose. Nothing had ever checked the gates.
METHOD, and it is the transferable part: for each gate, construct the label WITHOUT the fact - a
directory with the right name and no bake log, a handler case that exists only in a comment, a note
whose prose mentions the marker it lacks - run the gate, record what it says. No verdict was reached
by reading. Reading is how all five hid.
RESULT: 29 distinct scripts (35 registrations; three are shared across three runners). 19 sound, 4
holes left OPEN with rows, 6 that no plausible decoy could be built for and are named UNTESTED rather
than called sound. A gate nobody tried to fool is UNKNOWN.
SCOPE IS A FACT TOO - the largest single cause, and mundane. Eight gates decided what to look at with
os.listdir, one level. Every one was green AND CORRECT today, and every one would have gone blind the
moment anyone added a subdirectory. mojibake and docker-v already used os.walk, caught the identical
planted file, and are the control that proves the cause was the listing and not the decoy.
IN THIS REPO: hub-confirm and manifest-bearer now walk. observations_gate (R-419, CLOSED) requires a
marker at a line start or after a sentence boundary and strips inline code spans - a note SAYING it
carries no marker no longer satisfies the marker test. closed-register now CONVICTS on a row it
cannot parse instead of warning: FOUR rows were in that state, TWO of them written by the session
that closed them the day before, and every one was exempt from the only check that reads that file.
The rows were repaired first and the conviction added second - registering a failing gate refuses
every push.
THE META-GATE: decoy_coverage_gate.py refuses a gate registered without a decoy or a named exemption.
It convicted ITSELF the moment it was registered, which is how it came to have one. Coverage is a
DECLARATION the gate AST-parses, never a grep - searching a test file for a gate's name would be the
very shape this sweep exists to find. The 20 uncovered gates are listed by name (R-426).
NOT FIXED, each with a row and a decoy asserting TODAY's behaviour so the fix must be deliberate:
R-422 reuse-refs (only 7 extensions; a rotted .md citation is invisible), R-423 site (PAGES is a
hardcoded list of 7), R-424 one-register (a defect parked as `idea`), R-425 offbox-rename (fixed
FILES list). R-427: closed_register_gate checks ONE direction - twelve open rows carry a closed
verdict and were NOT moved, because telling finished from partly-finished is a judgement and R-378
is the record of a machine getting it wrong.
FIVE DECOYS WITHDRAWN AS ILLEGITIMATE, mine, named in the audit. A decoy nobody would write proves
nothing, and manufacturing a finding to fill a row is worse than an honest NO.
No product code. No version bump. No image. No golden owed. All four runners green.
Register: OPEN 172 -> 178, CLOSED 160 -> 161.