Vouch as a three-field change: golden 0.243.0 -> 0.244.0 with its new checksum;
agent and min_agent stay 0.131.0 because controller 0.244.0 declares the same
MinAgent. Floor raised 0.242.0 -> 0.244.0 with min_agent 0.131.0 so the hub does not
hold it — and it delivered: demo-felhom moved to 0.244.0 by itself within minutes.
VM 335 created from the BUILT 1.28.0 image, disks on /mnt/hdd_1, boot order set in
its own call. The boot menu proves the gate's menu criterion visually: two Hungarian
interactive entries and a 15 s countdown, no Proxmox entry, no automated entry.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
sha256 18328a3c7579628b8a7e9639777043db48063c86d37a2e0c221a6ccba6d755a0, 653 609 190
bytes, registry serves it. Markers: overlay2, both mount points, upload OK, no FATAL,
no publish-SKIPPED. Token-leak control passed with a planted positive control.
The three failures are written up because each is a rule this project already has:
scp -p instead of -P (nothing copied); the publisher run without the GITEA_USER it
requires, then the archive destroyed BEFORE checking the outcome; and a rewrite that
dropped the chmod, where the unit reported Result=success while the script inside it
had died on Permission denied.
The fix that matters is the gate: teardown now happens only when the REGISTRY serves
the package — not on an exit code, not on a log sentence. It held: on the failed
attempts the VM and its archive were left in place.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
It sits beside the 3-in-15-minutes budget in the host-agent design, because that
sentence and its measured blind spot belong together: four kills 20 minutes apart
were all restarted, none accumulated, and the only trace was an info event that
mails nobody. The budget itself is unchanged.
Also recorded: why Part C.2's re-issue button is correctly hidden for a customer
with no host, the venue's storage reconciliation (nvme-scratch IS /mnt/hdd_1), the
built ISO landing on demo-hp byte-identical, and my own scp/-P mistake that cost a
golden bake — including why its token-leak check reported a false hit on an empty
needle.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Closed with live proof on demo-hp (controller 0.244.0): the per-tier label and the
restore refusal. R-536 closed with its red-proofs and the hub's two new event types.
R-534 carries the measurement that matters: DatastorePowerUser is Backup+Prune only,
PBS has no custom roles, so DatastoreAdmin at the datastore root for the hub's user
is the narrowest grant that works. The row stays open until a re-issue is seen to
succeed end to end.
Opened: R-539 (a second, slower restart counter — the operator's ruling, for the
nightly), R-540 (one pool box, no selection rule when it fills), R-541 (no path to
move a customer between off-site boxes).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
ep0: DatastorePowerUser carries Backup+Prune only — measured, not recalled — so the
narrowest role that works is DatastoreAdmin, applied for the hub's user at the
datastore root only. Datastore.Modify now present; the per-customer DatastoreBackup
entries are untouched.
Tester 1's off-site tier is provisioned (shared, 100 GB) and a quota edit REUSES the
same sub-account (311327 all three times), 100 -> 150 -> 100 read back from the form.
R-537 and R-538 proven live on demo-hp running 0.244.0: Paperless's tier-1 row reads
„DB + Konfig" with the new sentence while tier 2 still reads „DB + Konfig + Adatok",
and pressing restore returns the Hungarian refusal with the app untouched.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Measured 2026-09-16: 25 minutes after a successful bind AND claim the console still
showed the pairing code under a line promising the screen refreshes itself.
print_bound_banner is printed the moment the bind delivery lands. It does NOT name
the dashboard URL: the one-shot delivery carries the customer id, passphrase and
mode, not the domain, so naming an address would mean inventing one. The residue —
the console still does not reflect the later CLAIM, because this unit has exited by
then — is recorded in the changelog rather than implied away.
Not published: the built image needs the release gate and the operator's yes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Off-site is ON by default for a new customer — shared, 100 GB soft quota prefilled,
the checkbox kept so an operator can opt a customer out. The reason is this repo's
own [FACT]: the whole-guest tiers do not carry the data drive and a Tier-1 unit has
no file leg, so with this unticked a one-drive box keeps NO copy of the household's
own files. Measured on a fresh box the same day.
The quota is prefilled because the fill warning only fires when quota_gb > 0.
Also registers controller v0.244.0's app_deploy_started / app_deploy_failed in both
allowedEventTypes and customerMessages, per the rule that the two move together.
Red-proofed: dropping the default fails the new render test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The automatic connect e-mail is proven with a real mailbox: the host record was
deleted at 12:22:59Z and the mail reached the customer at 12:23:00Z, one second
later, with selfbind_link_sent (host delete) on the timeline. The requirement was
two minutes. The hub refuses to delete an ONLINE host with no override, so the
record had to fall stale first — that wait is part of the proof.
Interventions: 0. Every P1 fix this drill set out to prove held on a fresh box.
The verdict is still no, for a new reason: a one-drive box with no off-site tier
keeps none of the household's own files in any backup, the page says otherwise,
and the restore that should save them makes it worse (R-537, R-538).
Teardown, three layers, stated. Customer tester-1 kept; RESET never used; nothing
on the off-site server written or removed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Machine and host layers are done and stated. The hub layer is deliberately waiting:
a host delete is refused while the host is ONLINE, with no override by design, so
the record must fall stale first — that wait is part of the proof.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Three more kills 20 minutes apart: recovered in 61 s / 41 s / 61 s, and none of
them accumulated, because the window is 15 minutes. Four restarts, zero pauses.
So the brake catches a FAST crash loop and is blind to a SLOW one — a controller
dying every 20 minutes is restarted forever, and the only trace is an info event
that mails nobody. Measured, not changed: the options are written into R-531 for
the operator to rule on.
Machine layer torn down: VM 334 purged with its disks, demo-hp's own containers
9201 and 9202 untouched. Evidence copied off the box first, token-leak control 0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Nothing on the walk needed a shell or an operator. The four moments that could be
mistaken for help are listed with the reason each is not one — two of them were my
own errors driving the API, and one was my own damage during the memory test.
The alarm table is now measured from two independent sides: the hub's own log lines
and the inbox. The one-hour operator cooldown is proven to suppress AND to release
(backup_tier_skipped mailed 12:08, suppressed 12:37 and 12:58, mailed again 13:18).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The off-site restore onto 9202 cannot be walked: this box never had an off-site
tier, because the re-issue fails on the endpoint token's missing Datastore.Modify
grant. Read-only listing of ep0 shows ns/tester-1/ct empty both before and after
the drill, with ns/demo-hp/ct as the positive control. Nothing on ep0 was written,
removed or pruned.
R-528 re-measured on a second, different box: all three OOM signals silent again.
R-511 records that its shipped fix is sound and inert until the grant is given.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
F10 (a child deletes the photo folder) is the finding: on a one-drive box with no
off-site tier the household's own files are in NO backup — the whole-guest tiers
exclude mp8 by design and the app's file leg lives at tier 2/3. The app page still
labels tier 1 „DB + Konfig + Adatok" (R-537), and the restore reports success while
leaving Nextcloud listing five photos it cannot open, after wiping the app's own
trash which still held every byte (R-538).
F11: the claim page locks out after the SECOND wrong code (15 minutes), the alarm
fires and is true; Nextcloud does not lock out. F12: two reboots 60 s apart, all
seven stacks back in 124 s, and the supervisor did not count the boots.
Evidence: audits/evidence-drill-0243-2026-09-16/phase2-f10.txt, -f11.txt, -f12.txt.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS