With the injector corrected, the hub really was unreachable. The controller
built its 23:08:42Z report, retried the push three times over 1m40.8s and
gave up at 23:10:23Z - 31 seconds before the link returned. Nothing queued,
which is correct: a report is a snapshot, not a fact.
The box passed. 26 containers throughout, every front door serving, both the
hub link and the host-agent link repaired unaided the moment the block lifted,
no alarm fired and none should have.
R-549 filed (P2): the staleness threshold (30 min) is exactly twice the report
cadence (15 min), so ONE failed push spends the entire budget. The measured gap
was 29m59s - one second inside the alarm. A healthy, self-repaired box came that
close to paging the operator.
Also recorded: the injected cut is broader than its name - it severed the
controller from its own host agent too, which a real ISP outage would not do.
The caveat travels with rounds 7, 8 and 9. The event-drop path remains
unmeasured, because no event was raised during any cut.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-547 (P3): a disk that fills and empties between sweeps is never mentioned to
anyone. The guest's root filesystem sat at 96% for ten minutes and no alarm of
any kind fired - checked twice, once by the round's runner and once
independently after the fill was released. disk_critical is defined at >=95%
used, but the fill-watch is a DAILY sweep plus one check ~90s after a controller
start, so a ten-minute window contains no check. The timing was almost comic:
the controller restarted at 21:28 after the previous round's power cut, so its
single opportunistic check ran about twenty seconds before the disk filled.
This is the ladder working as designed, not a missed alarm - it is filed because
the honest answer to "would the household be told?" is no, and that is written
down nowhere.
R-548 (P3): the whole-guest backup's LOCAL tier cannot fit on a
small-system-disk box and retries on that tier for ever. A ~29GB source into a
14GB pve-root, measured falling at ~16MB/s - under four minutes to a full / on
the nested PVE. The product's behaviour is correct throughout: it failed the
tier, named it, scheduled a retry, its status surface agreed, and the off-site
tier then succeeded from the same snapshot in ~8.5 minutes taking no local disk.
What is filed is the loop: on a box this shape the local tier can never succeed.
Honest caveat recorded in the row - the 32GB system disk is this drill's own
fixture choice - but nothing checks whether the local target could hold the
source before starting.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Round 2 (restore gokapi + power cut, 20s into the restore): the box came back
BY ITSELF in 148 seconds, 0 -> 25 -> 26 containers, and gokapi - the app being
restored when the plug came out - returned healthy. The only alarm was
controller_started, which is what the ladder expects for a 60-second outage:
no node_stale (30 min threshold), no app_start_failed (90s boot grace). No
false alarm, none missed.
Round 2's household measure is recorded as NOT COLLECTED, not as a pass: the
loop died with the box and zero lines is not zero failures.
The round also handed over immich's whole diagnosis. app_oom fired - "immich
(immich-postgres) - egy folyamatat a memoriakorlat leallitotta" - naming the
app and the exact container. That is why immich saw CONNECTION_CLOSED and
crash-looped twelve times. It is added as tonight's line on the EXISTING OOM
row rather than filed as a new one, because this project's standing finding is
that those signals are invisible inside LXC guests and on this box the scan
caught one. The diagnosis I spent twenty minutes reaching from logs was sitting
in the alarm feed, correctly labelled, the whole time.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The schedule was drawn from seed 20260917 and written into the findings document
BEFORE round 1, with its re-draw log.
Phase 0 measured:
- golden 0.245.0 baked, published (registry 200, not an exit code) and vouched;
the box installed itself from the published ISO 1.28.0 and landed on it with
no hand upgrade (controller 0.245.0, agent 0.131.0).
- ZERO operator presses: the waiting self-bind mail worked, and the acknowledged
-delete path re-issued off-site AND PBS-DR credentials by itself
(pbsdr_auto_reissue) - the F-14 half nobody had watched happen live.
- R-546 filed (P2): tonight's own guide sends the household to create the
recovery code ~17 minutes before the box can do it. It self-heals; the bar
urges them there the whole time. Measured on both sides, not inferred.
- R-543 proven through its whole lifecycle on a fresh box: bar present while
paused, gone for good once escrowed.
- Known rows met and recorded, not re-filed: R-542, R-536's failure events.
Also recorded honestly: three harness errors of mine (a script that announced
"all twelve deploys ACCEPTED" without checking, a "login ok (csrf 0)" that
turned eleven of my own 401s into what looked like product refusals, and a
head -12 that hid a disk), and a near-miss where I almost filed a defect
against a drive gate that was working and logging at DEBUG.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The tier-3 pause is the zero-knowledge escrow design and is untouched. What was
missing was the ASK, while the backup page promised the copy that had never run.
- VOLUNTEER-first-hour.md: a new step 6, right after the dashboard password and
before the first app - what the code is, where, write it on PAPER, and that
Felhom cannot get it back for them. Sections 6..12 renumbered to 7..13.
- day0-install.md A.2b: the operator step for a REBUILT box, which was missing.
Acknowledged delete -> the hub re-issues by itself; otherwise ONE press of
"Re-issue PBS credentials" (F-14 ruling 2026-07-13, hub/internal/web/pbsdr.go).
This is the correction to last night's "zero presses" note.
- 07-backup-architecture.md: 6.1 records tier-3's paused state as a DESIGN, and
2 records that the household is asked from first login.
- capability map: the first-hour row's last gap closed, with what it still does
not claim (no volunteer has walked the ask from the written guide).
- register: R-543 CLOSED with the live measurements; R-545 filed (nothing
un-configures an off-site target). R-511 was already closed yesterday.
- STATUS: the answered publish question removed (1.28.0 is live), readiness yes.
- evidence: red-proofs, the two-box live validation, teardown on three layers,
and both of my own mistakes in this session.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Uploaded with env-only credentials and verified by ROUND TRIP: the downloaded bytes
checksum to a4cd9b6d…, identical to the built file, and the checksum file is served.
1.27.1 stays in the bucket; nothing was overwritten.
The download page now names 1.28.0 with the published checksum (BOM preserved, site
gates green). R-535 closes with an honest caveat: the new banner ships byte-identical
to repo HEAD and the string is in the published payload, but it was never seen on a
screen — the box bound itself while the walk was headless.
Also corrected: the 1.27.1 heading still said NOT PUBLISHED although it went out on
the big night.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The capability map's journey row now carries the half it could never finish: five
photos in, deleted the way a child would, the old route refusing and touching
nothing, the off-site restore returning them, and them opening — sha256 identical,
5 of 5, with a negative control.
Stated with it, because both are true: the bind needed ZERO operator presses (the
box registered itself and used the mail the hub sent itself), but the PBS cascade
needed ONE — the Re-issue press R-511 documents, which then succeeded because of
this morning's ep0 grant.
R-543 (P1) is the honest caveat: off-site ON by default is not off-site WORKING on
day one — a fresh box waits at „Kulcsletétre vár" until the household creates its
recovery code, and nothing asks them to, while the tier-1 row already promises that
copy. R-544 records a log line that says „escrow deleted" where the effect is
demotion to retained custody.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
VM 335 purged with its disks; demo-hp's own containers untouched; evidence pulled
off the box before the destroy, with the one thing I could not collect stated (the
agent journal — root SSH is refused on the appliance by design).
The leak check redone properly: six real secret VALUES as needles against all 41
evidence files, planted positive control matched 6/6, committed evidence matched 0.
The earlier „22" was the word „password" in labels — a word count, not a leak check.
R-537 and R-538 now carry the fresh-box proof: the labels on a box where off-site is
on, the refusal that pointed at the off-site route, and five photos returned
byte-identical.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The rebuilt-customer case reproduced by itself: the WG-registration hook refused
exactly as R-511 describes and named the Re-issue action. Pressing it then worked —
reissue ok, pbsdr ADOPTED (gen 2), and the box consumed the single-use secret two
seconds later. No permission error.
This morning the identical action returned „missing Datastore.Modify … status 255"
and a 502. The only change in between is the narrow grant on ep0, and the narrowest
role was measured rather than recalled: DatastorePowerUser carries Backup+Prune only,
and PBS has no custom roles.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Measured on the fresh box: tier 3 sits at „Kulcsletétre vár" — the off-site copy is
paused until the household performs the key-escrow ceremony, and nothing asks them
to. POST /backup/offbox/run returns 302 and produces no snapshot; the controller log
shows only offsite-credential-retry.
That matters more after today, not less: the new default exists because a one-drive
box otherwise keeps the household's files in no tier at all, and the tier-1 row now
prints „Az alkalmazás fájljait a távoli másolat … védi". On day one that sentence
promises a copy that does not exist yet.
The good half, proven on the same page: tier 1 reads „DB + Konfig" with the new
sentence, and „DB + Konfig + Adatok" appears zero times — R-537 holds here too.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
I read „formatted but not mounted" off /api/disks/candidates and briefly held it as a
product fault. The storage page — the surface a household opens — says the opposite
and is right: Adatlemez, /mnt/felhom-drives/adatlemez, default, active, ext4.
R-542 records the real (small) defect: that endpoint offers a REGISTERED, in-use
drive under „initialize", with already_mounted null. The page filters it out, so no
customer sees it; it fed a formatting flow and it misled a session, which is enough.
Also recorded: off-site is LIVE on this fresh box by default — „Aktív — nincs
kijelölt alkalmazás" — an hour after that default shipped, with nobody pressing
anything.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Closed with live proof on demo-hp (controller 0.244.0): the per-tier label and the
restore refusal. R-536 closed with its red-proofs and the hub's two new event types.
R-534 carries the measurement that matters: DatastorePowerUser is Backup+Prune only,
PBS has no custom roles, so DatastoreAdmin at the datastore root for the hub's user
is the narrowest grant that works. The row stays open until a re-issue is seen to
succeed end to end.
Opened: R-539 (a second, slower restart counter — the operator's ruling, for the
nightly), R-540 (one pool box, no selection rule when it fills), R-541 (no path to
move a customer between off-site boxes).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Three more kills 20 minutes apart: recovered in 61 s / 41 s / 61 s, and none of
them accumulated, because the window is 15 minutes. Four restarts, zero pauses.
So the brake catches a FAST crash loop and is blind to a SLOW one — a controller
dying every 20 minutes is restarted forever, and the only trace is an info event
that mails nobody. Measured, not changed: the options are written into R-531 for
the operator to rule on.
Machine layer torn down: VM 334 purged with its disks, demo-hp's own containers
9201 and 9202 untouched. Evidence copied off the box first, token-leak control 0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The off-site restore onto 9202 cannot be walked: this box never had an off-site
tier, because the re-issue fails on the endpoint token's missing Datastore.Modify
grant. Read-only listing of ep0 shows ns/tester-1/ct empty both before and after
the drill, with ns/demo-hp/ct as the positive control. Nothing on ep0 was written,
removed or pruned.
R-528 re-measured on a second, different box: all three OOM signals silent again.
R-511 records that its shipped fix is sound and inert until the grant is given.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
F10 (a child deletes the photo folder) is the finding: on a one-drive box with no
off-site tier the household's own files are in NO backup — the whole-guest tiers
exclude mp8 by design and the app's file leg lives at tier 2/3. The app page still
labels tier 1 „DB + Konfig + Adatok" (R-537), and the restore reports success while
leaving Nextcloud listing five photos it cannot open, after wiping the app's own
trash which still held every byte (R-538).
F11: the claim page locks out after the SECOND wrong code (15 minutes), the alarm
fires and is true; Nextcloud does not lock out. F12: two reboots 60 s apart, all
seven stacks back in 124 s, and the supervisor did not count the boots.
Evidence: audits/evidence-drill-0243-2026-09-16/phase2-f10.txt, -f11.txt, -f12.txt.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
ISO 1.27.1 gated PASS and proven live: first-boot console Felhom-only,
pvebanner masked across a proven reboot. Hub v0.113.0 hand-over copy live
(R-497 closed). Full first hour walked again on customer tester-1 (three
disks + one disk): deploy, use, backup, removal, byte-identical restore,
power cut, typo all PASS. The tunnel gives a fresh box no routes: 12/12
503 from DooPlex (R-505); the record has no e-mail (R-508). Rows R-507,
R-508 filed; R-496/R-495 fixed/answered awaiting publish; day-0 A.1 no
longer claims the controller creates hostnames (R-506). NOT PUBLISHED.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Phase 0: the public ISO never auto-installs by construction (no answer.toml,
G1); the operator re-affirmed the interactive installer 2026-09-14.
- felhom-bootstrap.sh: mask pvebanner.service, write a Hungarian /etc/issue
(no :8006 admin URL); pairing banner names the Tulajdonosi jelmondat and
paints through the CONSOLE_DEV seam (R-496). Harness: 8 checks, red first;
fake hub now sends a pairing code (the banner was never tested, R-502).
- hub: created flash + Credentials block tell the operator to hand the phrase
over; the self-bind mail names the operator (R-497). Tests red first.
- iso-release-gate G14-G16; domain ruling in 01-topology + CONTEXT; R-494
narrowed to P3; R-502..R-504 filed; volunteer guide and day-0 A.2 aligned.
ISO_VERSION 1.27.0 (not built, not published).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Machine: VM 330 destroyed with its disks. Host: ISO and scratch removed,
~8.5 GiB returned on nvme-scratch. Hub: customer drill0242 DELETED via the
cascade (journal #17), pages 404; its ep0 WireGuard peer dropped at the next
full-list push. Evidence pulled before the destroy. R-501: the documented
CI-check recipe reads only the last jobs page, which is not in id order.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Golden 0.242.0 baked, round-trip verified and vouched (cadence rule, R-468).
Fresh box from the public ISO on demo-hp: landed on the vouched set, two apps
deployed and used, backup, remove, byte-identical restore, power cut and code
typo all PASS. Stopped for a volunteer by R-493 (no instructions) and R-494
(the setup mail's dashboard link has no DNS; intervention I1). R-493..R-500
filed. Capability map: first-hour row added (PARTIAL), journey row scoped.
Stopgap Hungarian volunteer guide written. Hub teardown layer pending.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Hub v0.112.0 serves a floor above the golden with a declared MinAgent;
controller v0.239.0 reached both demo boxes by that floor in 14 s and 15 s
and updates on any backup tier. 09 §3 decisions 7 and 8, §6/§6.1; 07 §6
line; capability map row; STATUS items 15/16 done and the cadence line
corrected; CONTEXT; register: R-470/R-472/R-475 compressed to CLOSED-ITEMS
(full text at 2f5d3af), R-477..R-480 opened, R-474 reproduced a third
time. OPEN-ITEMS 431689 -> 432156 bytes, CLOSED-ITEMS 118051 -> 120598.
Evidence: documentation/audits/rulings-r472-r475-2026-09-13/.
Controller v0.237.0-v0.238.1: the Update button is a guarded job — refusals, backup-first when the
proven Tier-2 copy is stale, safety dump, pin, pull (pin back on failure), health, HOLD on failure.
Proven live on demo-hp: A, B, E, F, H and the restore walk (audits/slice4-2026-09-13/).
Correction to this morning's pages: between golden bakes the hub HOLDS a floor above the vouched
golden, so a release does not reach the fleet by floor (R-472, operator decision). Corrected in the
runbook, STATUS, CONTEXT, R-468 and the gate docstring.
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS