- 09 §6.1 phase table (copying, undoing, undone), §6.1a SHIPPED with the two
live-only defects, §6.4 part 1 SHIPPED.
- Capability map: a failed update is undone by the box - PROVEN-LIVE.
- Live evidence on 9202: three apps undone by the product with seeds before
the backup, after it and seconds before the press read back; cut-off copy
held honestly; power cut during the undo resumed; manual press after undo.
- Register: R-637, R-639, R-641, R-642 closed; R-638, R-640 narrowed; R-643
ruled; R-646 opened. STATUS asks the floor question.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- Part 1 (09 §6.1a, audit): the undo performed by hand on 9202 for docmost
(PostgreSQL), romm (MariaDB) and vikunja (SQLite volume) - all three came
back with data written before AND after the backup. The product's loader
cannot do it: over a migrated PG database it fails on the new tables'
foreign keys; over MariaDB it leaves them behind. A truncated PG copy loads
rc 0 into an empty database. No-DB apps have no last-second copy.
- Part 2: one press jumps A -> C; the box's catalog clone is depth 1.
Ladder format recommended: update_ladder in .felhom.yml, not git history.
- Part 3: memory watch red-proof results (harness change in the catalog repo).
- Part 4 (09 §6.4): ten parts, ~22 evenings; one open point (R-643).
- Rows R-637..R-644 opened; R-446/450/451/462/463 updated. STATUS, CONTEXT.
No product code. Live catalog untouched; 9202 back on it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- 09 §3: decisions 11 (one window = a leg of the backup chain), 12 (automatic,
per-box switch on by default), 13 (the test decides, not the tag - replaces
decision 3's "never across a major"), 14 (the ladder), 15 (the box undoes a
failed update - replaces §6.1's no-auto-undo), 16 (Postgres majors converted
by the box), 17 (digests), 18 (fleet view, later).
- 09 §3b marked ANSWERED with a pointer per question; kept as the reasoning.
- 09 §6.2 rewritten to the ruled shape; §6.1 abort paragraph and §4 point at 15.
- Register: R-450, R-451, R-446, R-463 cite the decisions.
- STATUS: the seven questions no longer wait on the operator.
Documents only. No product code.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Both demo boxes took it in under 20 seconds and read healthy. drill-r50 is DOWN and correctly HELD:
its agent is 0.129.0, below the declared MinAgent, so the hub refuses to hand it a controller it
cannot run (R-472's guard, working).
Boxes below the floor: 5 -> 3. The three that remain are drill-r50 and the two guests the hub does
not hear from, including scratch 9202 which runs hub.enabled: false.
Read back from the hub rather than from the POST's own answer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
A (R-630, the P1): paperless-ngx, same app same button - failed at +313.0s with the app stopped
under v0.261.0, done at +53.4s now, with no "no probe container" warning because the explicit
healthcheck.container resolved the target.
B (R-633): the busy guard fires - RemoveStack REFUSED (busy): a backup or restore is running - and
the live proof caught it answering HTTP 500, because router.go maps remove errors by grepping the
error TEXT. v0.262.1 makes it a typed error and a 409; re-proven live.
C (R-634 half): a half-state with app.yaml on disk and deployed=false answered 200, leftovers NONE.
Under v0.261.0 the same call said "not deployed".
F (R-614): phase done before the remove, no phase at all after redeploying the same name.
Also: the first B run proved NOTHING and nearly went down as a pass - the refusal came from the
pre-existing "still running" check, not the new guard. Recorded.
09 6.1 and 8.8, the capability map, and STATUS updated. R-625 and R-634's mechanism are named as
owed, not half-done. Register 325.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The operator produced the mails. The controller HAS an OOM detector (main.go:821), it emits
app_oom (notifier.go:726), the hub allow-lists it (dispatcher.go:636) and delivered it to the
OPERATOR channel - two mails, 11:09 and 17:48 CEST, each naming the app and linking the dashboard.
CUSTOMER skipped, correctly.
I asserted an absence without opening the hub's Events or Notifications tab, reasoning instead from
a memory note that said the signal was UNPROVEN - not that it was missing. That is R-628's shape
again, from the same hand, four days later.
R-636: the real defect is the signal's SHAPE. notifier.go:715-724 keys on container|startedAt and
emits once per container lifetime, so 4530 worker kills over six hours produced exactly one
warning-level mail - indistinguishable from one transient kill. The magnitude was already collected
(App Telemetry: RomM 5023 errors, 632 warnings) but nothing turns it into a louder event.
Memory note lxc-docker-oom-signals-unreliable corrected with the positive reading.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found because the operator heard the fans. romm 5.3.0 was promoted that morning; the update read
`done` and the app ran clean for two hours, then OOM-crash-looped for six - 4530 worker SIGKILLs,
~500% CPU, host load 5.2 while otherwise idle, and nothing alarmed.
Raising the limit to 768M was still a guess and fixed nothing (memory.peak hit exactly 768 MiB).
Measured instead: ~216 MiB per warm uvicorn worker, so the image's default of 4 workers needs
~882 MiB. /init reads WEB_SERVER_CONCURRENCY; set to 2.
Proven under load, not just at idle: 26,645 requests over 300 s, memory 416-614 MiB against 768,
trending down, zero SIGKILLs, OOMKilled false. Idle CPU 500% -> 1.64%.
The first soak measured nothing - it was pointed at the scratch-guest subdomain, every request
404'd at traefik in 9 ms, and the counter reported 14,026 successes. Positive and negative controls
are now asserted before any load is driven.
Carried into R-462: `proven` has meant "the update applied and the data survived", not "the new
version runs".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
All 28 walked on scratch guest 9202 against the private drill catalog. 26 deployed, 6 proven,
5 inconclusive, 14 with no upstream edge, 1 failed honestly (outline 1.9.1->1.10.1, HELD with the
right sentence), 2 undeployable - one (plant-it) by design, refused by the lifecycle gate, proven
live for the first time. Each app also got the half the update night skipped: a restore from its own
copy with the seed read back again - 21 restored, 2 correctly REFUSED per 07 6.2.
R-630 RAISED TO P1 by measurement: a stack with NO probe container does not skip verifying - it
waits out the full health timeout and HOLDS, stopping an app whose three containers read healthy.
The controller's own words: "not healthy within 5m0s (last: no probe container)".
R-633 opened: a remove sent during a restore reports success and leaves a container restarting with
a live public route. The product already refuses that clash for update and for restore, naming the
blocker; remove has no such guard.
R-634 opened: an app can be running, healthy and serving while recorded as deployed=false, and is
then unremovable. Reproducible alone on sparkyfitness; concurrency-linked on two others.
R-631 and R-632 CLOSED. Register 321 -> 323. Seven interventions, six of them my own harness -
named, with what each cost. No product code. The live catalog's image: lines are byte-identical to
the start of the night.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Part 1: tandoor/zipline/wger probes corrected in the catalog and red-proofed live on 9202 in both
directions - "Nem egeszseges" with the front door serving 200, then "Fut" after the real sync with
no redeploy. tandoor's failed edge re-walked: done at +41.1s where it was failed at +361.9s.
Part 2: fifteen proven versions on the live catalog, one commit per app; the guarded Update pressed
on four apps on demo-hp, all four done.
Opened: R-630 (paperless-ngx's probe has never run on any box - a silent absence, worse than the
wrong probe that was found in one night), R-631 (five templates no static rule can judge),
R-632 (28 of 53 templates never deployed by any drill). Closed: R-618.
Register 318 -> 321. No product code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The drill is complete. Teardown done in three layers plus Gitea; the live catalog's every `image:`
line is proven identical to before.
WHAT WAS MEASURED. 21 edges across 19 apps, on scratch guest 9202 through the product's own
guarded Update, against a PRIVATE DRILL CATALOG so the live catalog carried no test reference at
any point: 14 proven, 3 failed, 4 inconclusive. Each app seeded and read back through its OWN
front door, with a negative control on every readback. Ten of the fourteen printed a verbatim
migration line. Up from the three apps this project had ever measured.
THE RESULT THAT MATTERS. R-618, P1: three of the 53 templates name a health probe the app does not
answer, and because the guarded update WAITS on that same probe, a SUCCESSFUL update ends by
STOPPING a working app. tandoor was measured serving HTTP 200 on the new version at four samples
across five minutes, docker's own healthcheck green, and was then stopped and the household sent
to a restore they did not need. zipline and wger are the same defect, both confirmed live. The
gate that catches all three is static and cheap: both health checks already sit in the same file.
WHAT THE NIGHT ANSWERED that was open. The UNATTENDED HOLD (312.9 s, pressed once, never again) —
which needed a purpose-built image store, because the rule that makes automatic updates safe is
the same rule that refuses the obvious way to break one. MariaDB across a major through the real
button, all four observables, first time. PostgreSQL across a major, refusing exactly as predicted,
with the conversion costed at ~9 s of engine work. There is NO single-flight: five updates ran at
once and all ended honest. And the two EARLY power-cut phases nobody had cut in.
TWELVE NEW ROWS (R-615..R-626), register 303 -> 315, and eight existing rows updated with what was
measured — including two CORRECTIONS: R-606 records the pre-flight refusals as reaching an English
household in English and they do not, and R-446/R-458 are both narrower than their rows state.
Two instrument fixes were needed before anything could be trusted: the unattended caller turned
every success into a timeout (R-623), and one of my own reproductions was wrong and is kept
labelled with what it actually measured.
Interventions: zero. No controller, agent or hub code written. The hub was never touched beyond
the floor the operator asked for.
Gates: repo_gates.py --fast, all 15 OK.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Part 0 — floor raised to 0.260.0, MinAgent 0.131.0 declared. 3 boxes below, all
down or blocked; both demo boxes SERVED.
Part 1 (R-610) — the DANGEROUS power cut, measured three times with three apps and
two cut mechanisms. All ended honest: resumed, completed, and pinned/installed/live
compose/docker inspect all agreed. vikunja's 2.6.0 migration had ALREADY run 0.64 s
after the cut decision and the seeded data read back intact — so the branch that is
one step from old-binary-on-migrated-database is now evidence, not argument.
Instrument limit stated: `starting` lasts under a second; all three landed in
`verifying`, which RecoverUpdates handles in the same branch.
Part 3 (R-611) — the night the previous session skipped without saying so. An app
updated with nobody pressing anything; a terminally-refused app was pressed exactly
once and never again over three passes. The unattended HOLD was NOT produced: the
within-a-major rule correctly refused the broken edge before it was attempted, so
Q4 still rests on the attended hold from slice 4. Said plainly rather than implied.
Rows: closed R-608/609/610/611; opened R-612 (P1 wishlist unusable on a fresh
install, and its error is a lie), R-613 (uptime-kuma healthy on its setup wizard),
R-614 (stale update phase survives a redeploy). R-520's pointer corrected.
Catalog: two drill pairs, both reverted; every image line byte-identical to
ff9717d3. The alpine:3.20 negative control a security review flagged is cleared.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Found by the live run, not by reading: POST /api/sync answered 'nincs valtozas'
while the box's catalog cache HAD moved, and catalog_images stayed stale until a
separate rescan. Since CatalogImages is the one input CatalogOrder compares
against, the badge answers from a stale catalog for that window — and the session
nearly recorded a stale tag-ok badge as proof of the R-524 ahead arm.
Neither half is isolated, so the row records the observation, not a diagnosis.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Phase 0 — measured, never estimated:
- both demo boxes: 10 apps, 0 behind, 0 unknown
- 46 of 58 exact catalog pins are behind upstream; 39 within a major, 7 across
- 6 of 7 measurable floating pins have been repushed since the catalog set them
(R-446 is no longer theoretical)
- the "23 of 66 floating pins" figure repeated in four places was STALE; recounted
to 10, with the definition written down beside it
Three claims in the brief corrected, named first:
- R-589 was NOT open — it shipped in v0.258.0; only the row was stale
- the chaos-night canary is NOT a defect — both gates refused to certify by design
- the hub half of the report confirmed, with the nuance that the raw payload is
stored whole, so Slice 7 is cheaper than the row implies
Closed: R-524 (controller v0.260.0, proven live in both languages), R-520 (power cut
during a REAL version change — the pin goes back, the app runs, the page says so),
R-589, R-469 (MariaDB half). Filed: R-605, R-606. R-462's stale scope corrected.
09 gains §3 decision 10 (decided by CC unattended — operator may reverse), §3b with
the seven questions in the decision shape, §6.2/6.3 the two open slices, and §6.4 an
update night costed from R-462's real numbers.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-601 said demo-hp was unreachable. The operator looked at the hub and said it
was online. It was, and had been up four and a half weeks, reporting every few
minutes. Both of my SSH routes pointed at stale addresses: `demo-hp` at a
tailnet peer for a box that has no tailscale installed at all, and `demo-hp-lan`
at 192.168.0.87 when the box is statically on .104 since a reprovision. The hub
had carried the right address in every host report, and `ip neigh` on felhom-pve
had .104 four lines above the .87 I quoted — I searched that output for the
address I expected instead of reading it for the address that was there.
Both ssh entries repointed and verified; nodes.md corrected, including that the
tailnet route for this box does not exist.
The hunt then found R-604, which is the real defect: demo-hp carried a
per-customer floor override of 0.243.0 left over from the 2026-09-16 drill, so
it had silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0.
`managed floor SERVED` fires once per change by design, so a box behind a static
override is silent for ever and its silence is indistinguishable from a box that
already logged. Cleared; demo-hp self-updated to 0.259.0 in under four minutes
and its claim page now answers "Wrong or expired code" in English.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The three blockers yesterday's English walk found are closed and each proven on
a live system. The verdict is deliberately 'nothing known now stands in their
way' rather than 'the walk passed': fixes are not a journey, and the hour has
not been re-walked by a stranger on a fresh install.
Also filed: the HP demo box answers on no route this session has (R-601), the
cookie-vs-session language instrument trap that would have had me fix R-598
twice (R-602), and the apostrophe that silently never matches a rendered page
(R-603).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Part 0 shipped as controller v0.258.0 (its own commit). Part A is the guide's
English twin. Part B is the walk: a box installed from scratch that day, used by an
English speaker following only the English guide and the screens.
VERDICT: not yet ready for an English-speaking tester, because the claim page — the
one screen between them and their box — is English chrome with Hungarian messages
(R-596, P1). Everything else held: the download page, the bilingual console, all
three customer mails, the bind page and its refusal, the dashboard's first language
from customer.language alone, both app pages, the whole catalog, the language switch
both ways — every one of them with zero Hungarian lines.
One intervention (I1 = R-494, filed 2026-09-14); the stop rule was not reached. The
walk exercised what 2026-09-14 could not: the graphical installer, the auto-reboot,
and the mailed link and self-bind page end to end — that walk's H1 is closed, because
this session had a mailbox.
R-214 CLOSED as a side effect and seen rather than reasoned about: the console's last
paint is now the bilingual "the box is linked" banner.
R-516 does NOT close, and the item-by-item note says why: more than half its twelve
items are about what a HUNGARIAN household reads, and an English walk cannot see
them. It now waits on a Hungarian walk with a second drive.
Rows opened: R-596 (P1, the claim page), R-597 (the setup code is three Hungarian
words), R-598 (the Backup page's protection warnings), R-599 (a drill's teardown is
blocked 30 minutes by report staleness and the 409 does not say so).
Golden 0.258.0 baked, published, vouched, with its record. The waiver was NOT retired
and the record says why in one line.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Part C shipped the same day the pilot was read: fifty apps in three pushes, 1 031 of
1 032 strings. The English Apps list shows ZERO Hungarian app descriptions across all
53 apps — the only Hungarian left on it is the "Naprakész" badge (R-589) and the
language picker naming itself, which is correct.
The Hungarian Apps list is identical to the pre-slice capture once the per-session
CSRF token AND Docker's own "Up N hours" container string are normalised. Both
normalisations are stated in the evidence rather than applied quietly — the second
one moved because two hours of wall clock passed between captures, not because any
copy changed.
Fleet floor raised to 0.257.0 with the declared MinAgent 0.131.0, above the vouched
golden so the declaration carries it. demo-felhom went 0.255.0 -> 0.257.0 by itself
in under 12 seconds and THEN rendered the English tagline: the floor delivered the
feature, not a version string.
Rows: R-593 (papra describes a session-signing key as "the app's subdomain" — the one
string left untranslated) and R-594 (the catalog gate can convict a retrieval promise
but has no way to REGISTER a true one, which the shared vocabulary's design calls
for). R-560 closed. 281 -> 287 rows.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
10-localisation.md §7 goes from [DESIGN, proposed] to [FACT], with the numbers it was
specced against corrected — and one correction chose an instrument rather than a
footnote. The catalog has 1 032 copy strings, not 835. 832 carry a Hungarian letter,
which was right. But the ASCII-ONLY Hungarian is ~120 strings, not three: „Aldomain"
appears 53 times and „A szerver domain neve" 53 times, and the three the plan named
(„Igen"/„Nem"/„Nincs") do not occur in this catalog at all. An accent-only gate passes
every one of them inside an English block — R-565's blind spot arriving again in a
different repo.
New §10.6 records what was proven live rather than reasoned about: the English pages
show English; the seven Hungarian pages are byte-identical before and after the push
apart from the per-session CSRF token; and a 0.255.0 box with the block synced onto it
renders identically and logs no warning, in a 93-line window that contains the sync's
own lines, so the absence is evidence and not a dead log.
Rows R-589 (the update badge is Hungarian on an English page), R-590 (the data-folder
card's backup promise, likewise, and it is a promise about the customer's files),
R-591 (Stack.Copy() deep-copies five Meta fields and not the new I18n map — safe
today, which is precisely why it is a row), R-592 (three defects inside the new
catalog gate, closed the same session, each found by its own decoy). R-560 updated:
Parts A and B done, Part C waiting on the operator's read of the pilot.
STATUS asks for that read, and for the floor to 0.257.0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Live at iso.felhom.eu, sha256 dceacae5da247d76cad065bf6c0d3bbefac8d8a5f8e571
db2fd8449a97e94829, and both download pages now name it.
Every gate criterion is recorded with its OBSERVED value in
documentation/tests/iso-release-1.29.0-2026-09-18/ — including two proof
installs from the published bytes, one per boot-menu entry, each with a first
boot AND one reboot: /etc/issue bilingual with zero hits for 8006, pvebanner
masked, package 1.29.0 installed, unit enabled and fired, pairing code present,
and the installed script byte-identical to repo HEAD. G11: the downloaded bytes
hash to the published checksum.
A defect was caught BETWEEN builds by looking at the screen rather than at the
config: the second menu entry read "Felhom telepítés (szöveges mód) / Install
Felhom (text mode)" — 58 characters — and the GRUB menu box cut it at "Instal".
The English half was unreadable on the boot screen. Shortened to "… / text" and
rebuilt; the published image is the rebuilt one. The Hungarian half is the part
that may not change, so the English half is the part that gave.
Teardown: VMs 323/324/325 destroyed, the two unclaimed appliance registrations
discarded (zero left in `registered`), guest 9201 untouched — 23 containers
before and after. The two stale *.rootpw.txt files were shredded from the
publish source directory before the upload ran from it (R-587, files gone; the
guard that would stop it recurring is still open).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Part B's live proof is a LOG LINE rather than an e-mail, and the evidence page
says why: severityNotifies drops `info`, so every safe converted producer mails
nobody, and every converted producer above `info` describes something bad that
is not true. Triggering one would mean a false record on a real box's timeline
or a state change the fences forbid. The log line was built in v0.256.1 for
exactly this, after finding there was nothing to look at on either side of the
wire.
15:07:28 controller_started [hu-only] — Controller elindult (0.256.1)
15:08:25 controller_started [+household(en)] — Controller elindult (0.256.1)
The Hungarian sentence is identical in both, and the hub stored the Hungarian
in every case including the English-household push.
R-558 CLOSED, R-555 closed with it. R-585 filed: six producers still send
Hungarian only because their sentence arrives already finished from another
package — `offbox_enlarge_blocked` matters most, since it has no hub entry so
its raw sentence IS the household's whole mail.
10-localisation.md gains 10.4; STATUS rewritten for the operator with the two
decisions left (raise the floor to 0.256.1; whether to rotate the demo
password after R-584).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
POST /configuration/global-floor with min_controller_version=0.255.0 and the
declared min_agent=0.131.0 (R-472); 303 flash=floor_set, read back from the
form, not from the POST.
The proof is the N100: it was never hand-deployed and its own Docker reports
felhom-controller:0.255.0 healthy within five minutes of the save. Three boxes
remain below — all BLOCKED or DOWN, which is a floor being held, not a floor
failing; each takes it on its next check-in.
R-580 filed: curl's %{redirect_url} rebuilds the request URL WITH the --netrc
credentials in it, so the hub password was printed into the session's own
output. Nothing written to a file, nothing committed. The build-deploy skill
now carries the rule.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-579 filed and closed the same day: five shell templates loaded style.css
with no cache-buster, so a browser holding the pre-0.254.0 file rendered the
new globe unstyled; and the globe sat outside the card.
- STATUS.md rewritten for the operator: what was seen, why, the third defect
found while fixing it (version disclosure on the guest share page, caught by
TestShareGuest_HeadersTilesNoAdminChrome), and the one decision left —
raise the fleet floor to 0.255.0, with what happens either way.
- 10-localisation.md §3: the shells' asset tag, and why the two guest pages get
an opaque tag rather than the version.
- Audit D: the parity diff (91 of 106 fixtures identical, every dashboard page
among them) and the live endpoint evidence from demo-hp guest 9201.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Agent versions checked FIRST (both live boxes 0.132.0, above the declared 0.131.0), because a
floor is held for a box whose agent is below the requirement. Positive observables at both ends,
and the box's is the one that counts: demo-felhom's own log reads "settle-gate: GO — at/above
floor 0.254.0", its image file and running container agree, and its four other containers stayed
up. The release's visible change is on that second box too: one globe on the sign-in page, zero
of the old text links.
What the table must not be read as: the two DOWN customers got nothing and will take 0.254.0
unattended when they next report, from 0.115.0 and 0.245.0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
10-localisation.md §10.3: the saved notes follow the box language at write time, with the
one-night consequence stated rather than hidden; the globe, and the table of WHO reads which
page and where its globe posts — getting that wrong makes the button do nothing, which it did
on /recovery until the live probe found it. Decision 6 superseded a second time; decision 8
(a claim carries the visitor's language) recorded. Decision 5 of §11's anonymous-surface line:
changing what a VISITOR reads is within what an anonymous request may do; changing anything the
household owns is not, and POST /lang can do only the first.
R-578 — the deadlock, and why it is a row rather than a fixed bug: UpdateOffboxStatus holds the
settings write lock while running its callback, boxLang() wants the read lock, sync.RWMutex is
not reentrant. On a real box an off-site run would have hung FOREVER holding that lock. The
symptom was a test suite going from 8 minutes to a 25-minute timeout. Fixed and guarded, but the
guard covers one package and three helper names; the class needs a gate.
R-577 — a guest share visitor still has no way to pick a language, and the household's setting
is the wrong default for a stranger. Deliberately left, pinned by a test, and the operator's to
decide because it is a promise the share feature makes.
.claude/rules/live-probes.md, unconditional: never send a deploy request for an app that is not
installed, not even expecting a refusal — the endpoint accepts first and validates later. Two
sessions made that mistake in two days, the second WITH a prompt line forbidding it. A prompt is
read once; a rule file is loaded every session.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The floor went 0.250.0 -> 0.253.0 with min_agent 0.131.0 declared, which is what carries a floor
above the vouched golden (R-472). Agent versions were checked FIRST, because the floor is held for
a box whose agent is below the requirement.
Positive observables at both ends, and the box's is the one that counts: the hub logs "managed
floor SERVED … from declared", and demo-felhom's own log reads "settle-gate: GO — at/above floor
0.253.0 (we are 0.253.0)" with the running image and four untouched containers to match. The
release then answered in both languages on that second box.
What the table must not be read as: the two DOWN customers got nothing and will take 0.253.0
unattended when they next report, from 0.115.0 and 0.245.0 — nobody has carried a box forward
from 0.115.0 in one step. Ordinary for a floor, and stated rather than left implied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
All 179 Hungarian error messages carry a key; zero remain. 10-localisation.md gains §10.2: the
four properties util.MsgError had to have at once and the failure each one prevents, and the
plural rule as a BUNDLE rule rather than a per-call-site flag, with the answerable sentence and
what the other option would have cost.
Two instrument defects recorded rather than tidied away, because both shapes recur:
R-576 — the parity gate has a measured blind spot. The bulk converter dropped the continuation
of multi-line concatenations, damaging 7 producers, and the gate stayed GREEN: every surviving
fragment WAS a byte-equal base-commit literal, so its question ("is this text real?") was
answered yes while the CALL had lost half its sentence. Two behaviour tests caught it. The
general form: a structural gate over the TEXT cannot see a defect in the CALL.
And the script counting what was left was case-sensitive, so it said "0 remain" while five did —
R-565's shape inside the measurement. Every "no Hungarian left" claim in this slice is now made
case-insensitively and with both controls.
R-575 — the soft memory-overcommit warning has no error to carry a key and no language where it
is built, so it renders Hungarian on an English page. Named in the code, not hidden.
Live evidence includes a mistake I made and corrected: a probe of the deploy refusal INSTALLED
vaultwarden on demo-hp (the endpoint accepts before it validates), the same mistake the previous
session recorded. Removed through the product's own path with its data; verified gone.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Endpoint-level on demo-hp guest 9201, method stated. The flash key renders the sentence in
each language where 0.251.0 showed the raw key; a link an OLD controller minted still shows
its prose, in both languages; country names follow the language and the list re-sorts.
Hungarian parity measured on the page BODIES, not on a hash: ten pages on 0.252.0, the guest
rolled back to 0.251.0, the same ten again, then rolled forward. Six byte-identical; the other
four differ only in live state (a clock crossing a minute, an app unhealthy for a moment, and
the update-available line, which is true on the old version and false on the new). The first
pass had saved only hashes — a hash cannot show WHAT moved — so the before-state was
reproduced rather than asserted.
Box left on 0.252.0 with its saved language `hu`; every English probe used the ?lang= override,
which is not persisted. Provisioned nothing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
- audits/r553-2026-09-17/: live before/after on demo-hp (CSRF redacted), the 409 refusal, the hub
health block, red-proofs for all five sites, the site-5 fixture diff, gates.
- 10-localisation.md §9: the five decisions with what each reads now; the rule (a text signature may
remain only where the text is not ours); the one legacy exception and its end date.
- Register: R-553 and R-563 closed to CLOSED-ITEMS; R-569 (four API handlers match English words),
R-570 (the legacy stale-note fallback + the slice-2 fence), R-571 (classifier and alert placement
documented nowhere). R-557 carries the R-570 dependency. 263 -> 264 open.
- STATUS, including the live probe that installed an app and was removed the same minute.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Design written after the spike ran (controller v0.247.0 live on demo-hp): mechanism,
flow, fallback, gates per language, catalog model, sliced plan with costs, operator
rulings 1-4 recorded, CC decisions 5-6, open decisions 1b and 7.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Decision 1 delivered: floor 0.246.0 served, the N100 on 0.246.0 within seconds;
Peti's box is DOWN on the hub and receives it when it reports.
Decision 2: the agent vouch was refused by R-120 until a newer golden existed.
On the operator's choice, golden 0.246.0 was baked (sha 05b7559d, amd64, all
markers, token leak 0 with control 1, registry 200 before teardown) and vouched
together with agent 0.132.0. golden_currency_gate: WAIVED -> OK. Bake evidence
filed where the gate and runbook read it: tests/golden-0.246.0-2026-09-17/.
Recorded, none reaching the registry: a first attempt on the arm64 template
(my version sort), a self-matching pkill, and an OOM-killed watcher whose
post-bake steps were done by hand.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-539 closed: five real controller kills on demo-hp 9201 with the production
24 h window raised controller_slow_crashloop, and exactly one operator mail
arrived (09:29:40Z). The fast brake never armed.
Register 215 -> 213 open (R-551, R-552 filed; R-539, R-546, R-549, R-550 closed).
unproven.py: 35 of 55 not walked, no number moved.
The report names the brief's wrong claims first and one recommendation not
followed: the controller floor was not raised - validated on one guest, a gap
of my own found during validation, and a floor above the golden reaches Peti's
box too. The operator's call.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
STATUS.md gets the morning note in the rules' order - decisions (none under the
unattended rule), what was exercised, what broke (nothing in the product; three
fixes worth making, all filed), rows (five opened, none closed), what could not
be tested, cleanup, and what needs the operator with the cost of doing nothing.
REPORT-chaos-night-2026-09-17.md follows template section 15 and names the
prompt's wrong claims first. It is a separate file because REPORT.md holds the
earlier session's write-up and this repo's rule forbids clobbering it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The tier-3 pause is the zero-knowledge escrow design and is untouched. What was
missing was the ASK, while the backup page promised the copy that had never run.
- VOLUNTEER-first-hour.md: a new step 6, right after the dashboard password and
before the first app - what the code is, where, write it on PAPER, and that
Felhom cannot get it back for them. Sections 6..12 renumbered to 7..13.
- day0-install.md A.2b: the operator step for a REBUILT box, which was missing.
Acknowledged delete -> the hub re-issues by itself; otherwise ONE press of
"Re-issue PBS credentials" (F-14 ruling 2026-07-13, hub/internal/web/pbsdr.go).
This is the correction to last night's "zero presses" note.
- 07-backup-architecture.md: 6.1 records tier-3's paused state as a DESIGN, and
2 records that the household is asked from first login.
- capability map: the first-hour row's last gap closed, with what it still does
not claim (no volunteer has walked the ask from the written guide).
- register: R-543 CLOSED with the live measurements; R-545 filed (nothing
un-configures an off-site target). R-511 was already closed yesterday.
- STATUS: the answered publish question removed (1.28.0 is live), readiness yes.
- evidence: red-proofs, the two-box live validation, teardown on three layers,
and both of my own mistakes in this session.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The capability map's journey row now carries the half it could never finish: five
photos in, deleted the way a child would, the old route refusing and touching
nothing, the off-site restore returning them, and them opening — sha256 identical,
5 of 5, with a negative control.
Stated with it, because both are true: the bind needed ZERO operator presses (the
box registered itself and used the mail the hub sent itself), but the PBS cascade
needed ONE — the Re-issue press R-511 documents, which then succeeded because of
this morning's ep0 grant.
R-543 (P1) is the honest caveat: off-site ON by default is not off-site WORKING on
day one — a fresh box waits at „Kulcsletétre vár" until the household creates its
recovery code, and nothing asks them to, while the tier-1 row already promises that
copy. R-544 records a log line that says „escrow deleted" where the effect is
demotion to retained custody.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The automatic connect e-mail is proven with a real mailbox: the host record was
deleted at 12:22:59Z and the mail reached the customer at 12:23:00Z, one second
later, with selfbind_link_sent (host delete) on the timeline. The requirement was
two minutes. The hub refuses to delete an ONLINE host with no override, so the
record had to fall stale first — that wait is part of the proof.
Interventions: 0. Every P1 fix this drill set out to prove held on a fresh box.
The verdict is still no, for a new reason: a one-drive box with no off-site tier
keeps none of the household's own files in any backup, the page says otherwise,
and the restore that should save them makes it worse (R-537, R-538).
Teardown, three layers, stated. Customer tester-1 kept; RESET never used; nothing
on the off-site server written or removed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
ISO 1.27.1 gated PASS and proven live: first-boot console Felhom-only,
pvebanner masked across a proven reboot. Hub v0.113.0 hand-over copy live
(R-497 closed). Full first hour walked again on customer tester-1 (three
disks + one disk): deploy, use, backup, removal, byte-identical restore,
power cut, typo all PASS. The tunnel gives a fresh box no routes: 12/12
503 from DooPlex (R-505); the record has no e-mail (R-508). Rows R-507,
R-508 filed; R-496/R-495 fixed/answered awaiting publish; day-0 A.1 no
longer claims the controller creates hostnames (R-506). NOT PUBLISHED.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Machine: VM 330 destroyed with its disks. Host: ISO and scratch removed,
~8.5 GiB returned on nvme-scratch. Hub: customer drill0242 DELETED via the
cascade (journal #17), pages 404; its ep0 WireGuard peer dropped at the next
full-list push. Evidence pulled before the destroy. R-501: the documented
CI-check recipe reads only the last jobs page, which is not in id order.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS