Commit Graph

1381 Commits

Author SHA1 Message Date
admin 3e6645413e CHAOS NIGHT: round 3 written up, and the household loop's blind spot stated
gates / gates (push) Successful in 21s
Round 3 in the findings document: the disk sat at 96% for ten minutes, twelve
apps kept serving, the household loop logged 12 operations with zero failures,
and NOTHING was ever raised. The silence is the finding, and it was predicted
from the ladder before the round: the fill-watch is a daily sweep plus one check
~90s after a controller start, and that single check ran about twenty seconds
before the disk filled.

Recorded with it: I twice labelled a mid-window reading "end of window",
estimating the clock instead of reading it. The readings were unchanged but the
label was wrong, and "nothing yet" is not "nothing ever".

And a limit of my own instrument, stated before its numbers get quoted: the
household loop does not follow redirects, so it measures "is the app serving on
the box" and never "can the household reach it from outside". It logged zero
failures straight through round 4's tunnel outage while the public route was
returning 530. So "0 household failures in round 4" must not be read as "the
household was unaffected" - someone away from home would have met 530 for about
ninety seconds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 23:46:02 +02:00
admin ec84eadc19 CHAOS NIGHT round 4: the box repairs its own tunnel in 97 seconds
gates / gates (push) Successful in 22s
Killed cloudflared and deliberately never restarted it, because whether it
returns by itself IS the measurement. It returned in ~97 seconds (killed
~21:41:30Z, running again 21:43:07.478Z), and the box did it, not Docker:
RestartCount=0 proves the unless-stopped policy never acted, and the controller
log shows "[infra] deploying cloudflared -> /opt/docker/stacks/cloudflared".
That is the product's protected-infra recovery repairing one of its own
infrastructure stacks unasked.

health_critical (error) fired at 21:43 - "Rendszer allapot kritikus (volt: ok)"
- which is exactly what the ladder predicts for a missing protected container,
and it is true. Whether health_recovered closes the pair is checked at the end
of the round, not guessed at now.

A METHOD CORRECTION that retro-labels every front-door reading tonight: the 530
during the outage is a Cloudflare status, which exposed that `curl -sL` was
following traefik's 301 out to the public hostname and back down the tunnel. So
every "front door" reading so far measured the PUBLIC path, not the LAN. It does
not invalidate the readings - a 200 by that route proves more, not less - but it
invalidates the label, and with it any claim of the form "the app is fine, only
the tunnel is down". The two paths are now measured separately: during this
outage the public route gave 530 while traefik answered 301 locally throughout.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 23:44:54 +02:00
admin 34d22a1a92 CHAOS NIGHT round 3: the disk fills for ten minutes and nobody is told
gates / gates (push) Successful in 21s
Round 3 (use bookstack, system disk held at 96% for ten minutes):
- the household saw nothing wrong: wiki, status and paste all answered 200
  before, during and after, and the background loop logged 12 operations with
  ZERO failures on a 96%-full disk
- the box kept all 26 containers running and released the space cleanly
  (29G used -> 944M used) with the thin pool untouched at 39.69% throughout
- NO alarm fired at any point, checked twice independently after the fill was
  released

That silence is the finding, and it was predicted from the ladder before the
round rather than discovered after: disk_critical is defined at >=95% used, but
the fill-watch is a daily sweep at 03:30 plus one check ~90s after a controller
start. The controller happened to restart at 21:28, so its single opportunistic
check ran about twenty seconds BEFORE the disk filled. A disk that fills and
empties between sweeps is invisible - by design, but the honest answer to
"would the household be told?" is no.

Also fixed and explained: my injector printed "unexpected EOF" while the
accident itself completed. bash -n passes, so it was not local syntax - G()
flattens its argument through `pct exec`, so a nested bash -c '...' has its
quoting re-parsed remotely. Both instances were in the disk branch only; the
accidents still to come use plain commands. The experiment was verified on the
box, not from the script's own account.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 23:42:18 +02:00
admin aca0172efd CHAOS NIGHT: fences clean, firewall baseline taken, and a third mistimed label
gates / gates (push) Successful in 22s
Fence check mid-night: demo-hp carries only this drill's VM 336; guests 9201 and
9202 are running with their full standing sets intact, bentopdf included.
Nothing of theirs was stopped, removed or redeployed tonight.

Firewall baseline recorded BEFORE the rounds that need it (7-9 block the box's
internet by flipping a host sysctl and inserting two physdev rules):
  iptables -S FORWARD -> "-P FORWARD ACCEPT" and nothing else
  physdev rules -> 0
  net.bridge.bridge-nf-call-iptables = 0
A control taken before the experiment, so that "it looks clean afterwards" can
be a measurement rather than an assertion - on a host that also carries the two
standing demo guests.

And a third mistimed reading of my own, recorded: I labelled a 21:36:00Z check
"end of window" when the fill runs to ~21:39:49Z. The readings are unchanged
(nothing fired), but the label is the point - "nothing yet, five minutes in" and
"nothing in the whole window" are different findings. From here the end-of-window
check is taken when the round's runner reports completion, because the runner
knows when it released the fill and I was guessing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 23:36:41 +02:00
admin ee3da86d33 CHAOS NIGHT round 3: interim reading, and a mislabel caught before it stood
gates / gates (push) Successful in 20s
Five minutes into the 96%-full disk: the apps keep serving (26 containers), the
pool is untouched at 39.69%, the filesystem is writable, and nothing has been
raised - the newest event is still controller_started from 21:28.

I nearly filed this as the "end of window" check. The fill began 21:29:49Z, so
the ten-minute hold runs to ~21:39:49Z and this reading was taken at 21:34:23Z,
halfway through. It is recorded as INTERIM and the end-of-window check stays
owed, because an alarm arriving late is a different finding from one that never
arrives - and a mid-window reading standing in for the final one would have
quietly turned "not yet" into "never".

What it already establishes: twelve apps keep running and serving with the
system disk at 96% full, and after five minutes nobody has been told anything.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 23:35:13 +02:00
admin fa1ddd92a5 CHAOS NIGHT: the internet-block accident now cleans up unconditionally
gates / gates (push) Successful in 20s
Reviewed before it runs unattended at ~01:11, not after. The accident flips a
host-wide sysctl on demo-hp and inserts two FORWARD rules, and its cleanup ran
only on the happy path: if the script were killed during its ten-minute sleep,
or the SSH dropped, the rules and the sysctl would have stayed. demo-hp is a
Tier-0 host that guests 9201 and 9202 also live on, so an abandoned FORWARD
rule is a fence breach rather than a measurement.

It now traps EXIT, INT and TERM, removes both rules and restores the sysctl
whatever happens, and clears the trap on the normal path so the cleanup does
not run twice. The rules still match --physdev-in on this VM's own tap, resolved
at run time, and the default FORWARD policy is never touched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 23:34:00 +02:00
admin 7221367f38 CHAOS NIGHT round 3: the disk fills, and nothing is told about it
gates / gates (push) Successful in 21s
Measured while the guest's root filesystem was held at 96%:
- the app's view is a genuinely full disk (29G used, 1.5G free, 28G fill file)
- the shared thin pool stayed at 39.69% - fallocate reserves blocks without
  writing them, so this round is NOT a repeat of the pool exhaustion that
  wedged the box in Phase 0. Recorded explicitly, because "disk 95% full"
  invites exactly that wrong reading.
- the filesystem stayed writable (a real touch, not the mount flags)

No disk alarm fired, and that was PREDICTED from the ladder before the round:
the fill-watch is a daily sweep at 03:30 plus one check ~90s after a controller
start. The timing is sharper still - the controller restarted at 21:28 after
round 2's power cut, so its single opportunistic check ran about twenty seconds
BEFORE the disk filled. A disk that fills and empties between checks is
invisible; that is by design, but it is the honest answer to "would the
household be told?" - no.

Household lines are attributed to the right round: the two UNREACHABLE entries
at 21:27:57Z are round 2's recovery tail, not round 3's accident.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 23:33:08 +02:00
admin bca013edec CHAOS NIGHT: round 2 passed, and the OOM finding goes on the row that owns it
gates / gates (push) Successful in 23s
Round 2 (restore gokapi + power cut, 20s into the restore): the box came back
BY ITSELF in 148 seconds, 0 -> 25 -> 26 containers, and gokapi - the app being
restored when the plug came out - returned healthy. The only alarm was
controller_started, which is what the ladder expects for a 60-second outage:
no node_stale (30 min threshold), no app_start_failed (90s boot grace). No
false alarm, none missed.

Round 2's household measure is recorded as NOT COLLECTED, not as a pass: the
loop died with the box and zero lines is not zero failures.

The round also handed over immich's whole diagnosis. app_oom fired - "immich
(immich-postgres) - egy folyamatat a memoriakorlat leallitotta" - naming the
app and the exact container. That is why immich saw CONNECTION_CLOSED and
crash-looped twelve times. It is added as tonight's line on the EXISTING OOM
row rather than filed as a new one, because this project's standing finding is
that those signals are invisible inside LXC guests and on this box the scan
caught one. The diagnosis I spent twenty minutes reaching from logs was sitting
in the alarm feed, correctly labelled, the whole time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 23:31:13 +02:00
admin 5b6e4b5c30 CHAOS NIGHT: round 2 under way, and the household loop made to survive accidents
gates / gates (push) Successful in 21s
Round 2 (restore gokapi + power cut) is running: the restore started, the plug
came out 20s later, and the box is recovering on its own.

Found at round 2 and fixed for the rest of the night: the background household
loop ran as a TRANSIENT unit on the VM, so the first accident that could have
produced household failures - a power cut - instead killed the loop and produced
no lines at all. Zero lines is not zero failures, and round 2's household
measure is recorded as NOT COLLECTED rather than as a pass. It is now a real
systemd unit with Restart=always, enabled at boot, so it returns with the box
after the power cuts, hard reset and docker restart still to come.

Machinery for the remaining rounds written in advance rather than mid-round:
one generic runner covering every action and accident the seed actually drew,
so no round is measured a different way from another. It records the same five
things each time, counts the household loop's lines and failures for its own
window, and names immich as a known pre-existing failure so nothing later is
misattributed to an accident.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 23:29:10 +02:00
admin cc87efa235 CHAOS NIGHT: household whole (11/12), injector fixed, round 2 running
gates / gates (push) Successful in 21s
The household the accidents actually hit: 11 of 12 front doors serve 200,
verified with a no-such-host control so a 200 means a real route to a real app.

bookstack's 500 was my seventh error and the first I confirmed before acting:
the repair script DECODED the stored APP_KEY and printed "prefix ok: False,
body length: 96, NOT valid base64" - Laravel could never have used it. A clean
redeploy with base64:$(openssl rand -base64 32) had it healthy in 45 seconds.

immich is diagnosed (CONNECTION_CLOSED to its postgres during reverse-geocoding
init, RestartCount=12) and deliberately LEFT BROKEN: no round in the drawn
schedule acts on it, and chasing the one app the schedule never touches would
cost rounds that were drawn before the night began. It is named as a known
pre-existing condition so no later failure is misattributed to an accident.

Machinery fixed before its round arrives, not during it:
- inject.sh used `qm guest exec`, which this box cannot do (no guest agent,
  measured in Phase 0). Three drawn accidents depend on in-guest work, so it now
  goes over SSH + pct exec.
- the disk-95%-full accident gained a POOL GUARD: the guest's disks are thin
  provisioned over the pool that hit 100% and remounted the box read-only
  earlier tonight, so the fill is capped and any cap is declared in the round's
  own evidence rather than silently applied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 23:26:03 +02:00
admin a046db7df7 CHAOS NIGHT: household verified, and a seventh error of mine found by control
gates / gates (push) Successful in 21s
Verified after the repairs: 26 containers running, all twelve apps deployed,
ZERO auth-failure lines on the rebuilt DB-backed apps, and 10 of 12 front doors
answering 200 - including cloud (nextcloud) and share (gokapi), both of which
were broken an hour ago. The cures are confirmed at the front door, not by a
health badge.

bookstack diagnosed properly rather than guessed at: its healthcheck exits 22
(curl's "server returned an HTTP error"), a direct request to the container
returns 500, its migrations completed cleanly and it has no auth failures. So
neither the database nor the image is at fault - the app itself errors. The
likely cause is mine: I passed APP_KEY=base64: plus 32 random alphanumerics,
which is not a base64-encoded 32-byte key. The repair decodes the stored key
and prints its true length BEFORE redeploying, so the hypothesis is confirmed
or refuted in the evidence.

Also recorded: my sixth slip, running docker over SSH on the VM instead of
inside the guest, which printed a tidy table of "absent" and "0" that read like
"nothing is wrong" and was produced by a shell with no docker at all. Six of my
errors tonight share one shape - a command whose precondition failed, still
printing a confident answer - and the same discipline caught every one: ask the
box directly, with a control.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 23:23:04 +02:00
admin c3e1986aaf CHAOS NIGHT: household repaired, headroom checked, round 2 armed
gates / gates (push) Successful in 21s
The five apps my disk-full burst broke are removed with their data and
redeployed cleanly, one at a time. Two distinct faults of mine, with different
cures, and separating them is what made either fixable:
  - image layers written while the pool was full -> "invalid ELF header",
    exit 127; cured by dropping the image so compose re-pulls
  - my re-seed's FRESH database passwords over volumes initialised with the
    first set -> Postgres auth_failed / MariaDB "Access denied"; cured only by
    removing the app with its data and deploying once

gokapi proves they are different: a new image left it Restarting(1), a new
database made it healthy.

Headroom measured so the disk-full failure cannot quietly repeat: pool 39% of
75.8G, docker filesystem 29%, data drive 1%, 4.3G guest RAM free.

Round 2's script hardened before it runs unattended: its "steady" test now
compares against the count it measured itself in the same round, instead of a
hardcoded 24 that the rebuilt household might never reach.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 23:20:27 +02:00
admin a1a57ea6d5 CHAOS NIGHT: round 1 recorded, and three of my own conclusions corrected
gates / gates (push) Successful in 20s
Round 1 (offsite-run, control round) is written up with its five things: the
run worked end to end in 1m45s, and on this REBUILT box the remote repository
is orphaned - the documented behaviour, surfaced honestly instead of reported
as a successful copy. Three alarms fired, all true and precise.

Corrections to my own earlier claims, each recorded where the wrong version
was written:
- "the 404s were my mistimed sweep" - wrong for four of five. Proven with a
  negative control (a no-such-host request returns the identical 404, 19 bytes)
  that traefik simply has no route to an unhealthy container.
- "nextcloud is repaired" - wrong. The re-pull fixed the corrupt library, but
  the app still cannot reach its database, and the container reports HEALTHY
  the whole time. A health signal is not a data signal.
- "all the broken apps are corrupt layers" - wrong. bookstack logged a clean
  startup, gokapi logged nothing, and immich shows a Postgres auth_failed.

The real cause of most of it is mine: my re-seed generated FRESH database
passwords over volumes whose databases were initialised with the first set.
The affected apps are being removed with their data and redeployed cleanly.

Also recorded: an HTTP 000 is "no answer", not "it failed" - the action still
took effect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 23:17:39 +02:00
admin 5f2ccec64c CHAOS NIGHT: household seeded, escrow done, round 1 measured
gates / gates (push) Successful in 19s
Phase 0 finished: all twelve apps deployed one at a time (the parallel burst was
my error and is recorded as such), escrow ceremony completed after the box
converged the PBS descriptor by itself (~17 min), and the R-543 bar disappeared
for good once escrowed.

Round 1 (offsite-run, control round, no accident):
- the run started on the button's own endpoint and walked every app with real
  per-volume byte counts (stop -> dump -> restart)
- three alarms fired, all TRUE and precise: app_start_failed named the one
  crash-looping app, backup_run_failures said "1 of 12 ... nextcloud", and
  offbox_repo_orphaned reported the documented rebuild behaviour rather than
  claiming a successful copy
- nextcloud's crash loop is MY damage (image layers written while the thin pool
  was 100% full -> "invalid ELF header"), not a product defect, and is recorded
  that way

The catalog bump was prepared and then REVERTED UNPUSHED: the catalog repo's own
gates returned INCONCLUSIVE (the volume-persistence prober failed its own
canary), and undetermined is never a pass. Round 7's `update` therefore becomes
`use`, decided now rather than improvised at 02:00.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 23:10:37 +02:00
admin 9fae6dfa98 CHAOS NIGHT phase 0: golden 0.245.0, a self-installing box, and R-546
gates / gates (push) Successful in 23s
The schedule was drawn from seed 20260917 and written into the findings document
BEFORE round 1, with its re-draw log.

Phase 0 measured:
- golden 0.245.0 baked, published (registry 200, not an exit code) and vouched;
  the box installed itself from the published ISO 1.28.0 and landed on it with
  no hand upgrade (controller 0.245.0, agent 0.131.0).
- ZERO operator presses: the waiting self-bind mail worked, and the acknowledged
  -delete path re-issued off-site AND PBS-DR credentials by itself
  (pbsdr_auto_reissue) - the F-14 half nobody had watched happen live.
- R-546 filed (P2): tonight's own guide sends the household to create the
  recovery code ~17 minutes before the box can do it. It self-heals; the bar
  urges them there the whole time. Measured on both sides, not inferred.
- R-543 proven through its whole lifecycle on a fresh box: bar present while
  paused, gone for good once escrowed.
- Known rows met and recorded, not re-filed: R-542, R-536's failure events.

Also recorded honestly: three harness errors of mine (a script that announced
"all twelve deploys ACCEPTED" without checking, a "login ok (csrf 0)" that
turned eleven of my own 401s into what looked like product refusals, and a
head -12 that hid a disk), and a near-miss where I almost filed a defect
against a drive gate that was working and logging at DEBUG.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 22:56:08 +02:00
admin d124c77e17 R-543 closed: the household is asked for the recovery code (controller v0.245.0)
gates / gates (push) Successful in 21s
The tier-3 pause is the zero-knowledge escrow design and is untouched. What was
missing was the ASK, while the backup page promised the copy that had never run.

- VOLUNTEER-first-hour.md: a new step 6, right after the dashboard password and
  before the first app - what the code is, where, write it on PAPER, and that
  Felhom cannot get it back for them. Sections 6..12 renumbered to 7..13.
- day0-install.md A.2b: the operator step for a REBUILT box, which was missing.
  Acknowledged delete -> the hub re-issues by itself; otherwise ONE press of
  "Re-issue PBS credentials" (F-14 ruling 2026-07-13, hub/internal/web/pbsdr.go).
  This is the correction to last night's "zero presses" note.
- 07-backup-architecture.md: 6.1 records tier-3's paused state as a DESIGN, and
  2 records that the household is asked from first login.
- capability map: the first-hour row's last gap closed, with what it still does
  not claim (no volunteer has walked the ask from the written guide).
- register: R-543 CLOSED with the live measurements; R-545 filed (nothing
  un-configures an off-site target). R-511 was already closed yesterday.
- STATUS: the answered publish question removed (1.28.0 is live), readiness yes.
- evidence: red-proofs, the two-box live validation, teardown on three layers,
  and both of my own mistakes in this session.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 21:20:51 +02:00
admin 1acd693854 the volunteer's own path verified: felhom.eu/letoltes serves 1.28.0
gates / gates (push) Successful in 21s
The page names 1.28.0 four times and 1.27.1 zero times, the link answers 206, and the
published checksum is the checksum of the file that was gate-checked and installed
tonight. CI green for both repos, matched by head_sha (669, 648).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 20:27:25 +02:00
admin 832218dca4 ISO 1.28.0 PUBLISHED on the operator's yes; download page and R-535 updated
gates / gates (push) Successful in 20s
Uploaded with env-only credentials and verified by ROUND TRIP: the downloaded bytes
checksum to a4cd9b6d…, identical to the built file, and the checksum file is served.
1.27.1 stays in the bucket; nothing was overwritten.

The download page now names 1.28.0 with the published checksum (BOM preserved, site
gates green). R-535 closes with an honest caveat: the new banner ships byte-identical
to repo HEAD and the string is in the published payload, but it was never seen on a
screen — the box bound itself while the walk was headless.

Also corrected: the 1.27.1 heading still said NOT PUBLISHED although it went out on
the big night.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 20:25:55 +02:00
admin 3f7ac8ee6e the backup promise is kept: photos deleted and returned byte-identical
gates / gates (push) Successful in 20s
The capability map's journey row now carries the half it could never finish: five
photos in, deleted the way a child would, the old route refusing and touching
nothing, the off-site restore returning them, and them opening — sha256 identical,
5 of 5, with a negative control.

Stated with it, because both are true: the bind needed ZERO operator presses (the
box registered itself and used the mail the hub sent itself), but the PBS cascade
needed ONE — the Re-issue press R-511 documents, which then succeeded because of
this morning's ep0 grant.

R-543 (P1) is the honest caveat: off-site ON by default is not off-site WORKING on
day one — a fresh box waits at „Kulcsletétre vár" until the household creates its
recovery code, and nothing asks them to, while the tier-1 row already promises that
copy. R-544 records a log line that says „escrow deleted" where the effect is
demotion to retained custody.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 20:24:50 +02:00
admin c18efc0610 teardown layer 1, a proper secret-leak check, and the fresh-box proof in both rows
gates / gates (push) Successful in 20s
VM 335 purged with its disks; demo-hp's own containers untouched; evidence pulled
off the box before the destroy, with the one thing I could not collect stated (the
agent journal — root SSH is refused on the appliance by design).

The leak check redone properly: six real secret VALUES as needles against all 41
evidence files, planted positive control matched 6/6, committed evidence matched 0.
The earlier „22" was the word „password" in labels — a word count, not a leak check.

R-537 and R-538 now carry the fresh-box proof: the labels on a box where off-site is
on, the refusal that pointed at the off-site route, and five photos returned
byte-identical.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 19:51:54 +02:00
admin c91c1377bc THE PHOTOS OPEN — the backup promise is true on a box that installed itself today
gates / gates (push) Successful in 22s
Five photos in, deleted the way a child would, and back again: HTTP 200 with bytes
200000/400000/600000/800000/1000000 and sha256 identical to the originals, five of
five, with a negative control. Controls at the same moment: status.php 200, WebDAV
207.

The two halves that make it honest:
- the OLD route refused and touched nothing („a fájlok így a helyükön maradnak"),
  naming the route that could help; the app was running before and after;
- the off-site restore ran in two steps — a verification copy that states „A meglévő
  adatok változatlanok", then a reconstitution whose message counts „5 fájl és 3
  adatkötet és az adatbázis".

This morning the same deletion ended with five photos listed, none of them openable,
and a success message over the top.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 19:50:05 +02:00
admin c5df1a372b the off-site tier works on the fresh box, and the photos are deleted for the test
gates / gates (push) Successful in 20s
The escrow ceremony ran (start re-authenticates, phase done, uploaded, sealed) and
the one-time recovery code was captured into a 0600 file at the moment it appeared —
it is shown once, the box stores it nowhere, and it appears in no committed file.

Tier 3 then reported „Sikeres restic → …your-storagebox.de · Helyreállítási egység,
titkosítva", and „Kulcsletétre vár" is gone. The restore wizard renders for Nextcloud,
which it only does for an app the store can actually restore — checked BEFORE
deleting anything, because deleting with an unproven copy is the harm itself.

Then the child's action: DELETE /Fotok -> 204, the folder 404s, the photo 404s.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 19:45:17 +02:00
admin a73abf04db R-534 and R-511 CLOSED — the ep0 grant proven end to end on a fresh box
gates / gates (push) Successful in 20s
The rebuilt-customer case reproduced by itself: the WG-registration hook refused
exactly as R-511 describes and named the Re-issue action. Pressing it then worked —
reissue ok, pbsdr ADOPTED (gen 2), and the box consumed the single-use secret two
seconds later. No permission error.

This morning the identical action returned „missing Datastore.Modify … status 255"
and a 502. The only change in between is the narrow grant on ep0, and the narrowest
role was measured rather than recalled: DatastorePowerUser carries Backup+Prune only,
and PBS has no custom roles.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 19:26:54 +02:00
admin 63e2de9b9e R-543: off-site ON by default is not off-site WORKING on day one
gates / gates (push) Successful in 21s
Measured on the fresh box: tier 3 sits at „Kulcsletétre vár" — the off-site copy is
paused until the household performs the key-escrow ceremony, and nothing asks them
to. POST /backup/offbox/run returns 302 and produces no snapshot; the controller log
shows only offsite-credential-retry.

That matters more after today, not less: the new default exists because a one-drive
box otherwise keeps the household's files in no tier at all, and the tier-1 row now
prints „Az alkalmazás fájljait a távoli másolat … védi". On day one that sentence
promises a copy that does not exist yet.

The good half, proven on the same page: tier 1 reads „DB + Konfig" with the new
sentence, and „DB + Konfig + Adatok" appears zero times — R-537 holds here too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 19:21:28 +02:00
admin 0cfbfe709c R-536 proven live on a fresh box: acceptance no longer claims the app is installed
gates / gates (push) Successful in 20s
At the moment of the 202 the hub received app_deploy_started — Alkalmazás telepítése
elindult: Nextcloud — and app_deployed is absent while the install is still running.
This morning the same moment produced „Alkalmazás telepítve: Mealie" for an install
that was killed five seconds later and never happened.

Also recorded: the deploy was first refused 400 because the admin password field is
mandatory at the server while its own metadata says required:false with
generate:password:16 — the browser fills it with the Generálás button, so a household
never meets it, but an API caller that trusts the metadata does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 19:16:39 +02:00
admin 2687819195 the drive was fine and my instrument was not — corrected, and R-542 filed
gates / gates (push) Successful in 20s
I read „formatted but not mounted" off /api/disks/candidates and briefly held it as a
product fault. The storage page — the surface a household opens — says the opposite
and is right: Adatlemez, /mnt/felhom-drives/adatlemez, default, active, ext4.

R-542 records the real (small) defect: that endpoint offers a REGISTERED, in-use
drive under „initialize", with already_mounted null. The page filters it out, so no
customer sees it; it fed a formatting flow and it misled a session, which is enough.

Also recorded: off-site is LIVE on this fresh box by default — „Aktív — nincs
kijelölt alkalmazás" — an hour after that default shipped, with nobody pressing
anything.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 19:04:20 +02:00
admin 4209908415 fresh box claimed, and it landed on the golden this task published
gates / gates (push) Successful in 21s
Login with the new password returns 302 and the dashboard opens: the claim took.
The page reads controller 0.244.0 — so a box installed from the built ISO lands on
the vouched set with no hand upgrade (agent 0.131.0, controller 0.244.0, PBS wrapper
matching).

Recorded alongside: a transport fact (the same claim page 403s to a python client and
200s to curl seconds apart — the edge judging the client, not the box refusing), and
that the drive init's {"started":true} is an attempt, not a result, so nothing is
deployed onto the drive until the mount is observed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 18:20:05 +02:00
admin a1e9ff6771 the fresh box enrolled itself and the claim mail arrived — both unprompted
gates / gates (push) Successful in 20s
Host tester-1-33b6a9 is ONLINE minutes after the bind, on the vouched agent 0.131.0
with the PBS wrapper matching the vouched hash, and its capability list already
reports the felhom-pbs backup tier readable by the agent. The customer guest was
still being created at that moment.

The setup-code mail arrived by itself at 16:00:58Z. The code is a secret: it is held
out-of-band for the claim step and appears in no committed file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 18:02:00 +02:00
admin 8f5b504c5c fresh box installed, first boot is Felhom's, bound with ZERO operator presses
gates / gates (push) Successful in 21s
First boot of the installed system shows Felhom's own Hungarian screen, pairing code
ZB3-7HM, and no Proxmox admin URL (8006 appears zero times). The box registered
itself on the hub from the universal secret-free image — same code, same MAC — with
nothing pressed on the operator side.

The bind then used the mail the hub sent ITSELF after this morning's host delete
(R-509), so the operator press the previous drill needed is gone: „Sikeres
összekötés." The form's field names were read, not guessed.

Also recorded: the reboot trap reproduced exactly as documented (a completed install
looks identical to a stuck one, so completion was judged from behaviour); and the
owner passphrase was handled file-to-file, which is the correction to this project's
one real secret slip.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 18:00:52 +02:00
admin 70f491eb8b fresh box: the built ISO 1.28.0 installs — every screen read before every keystroke
gates / gates (push) Successful in 20s
Twenty-one screendumps from the machine's own console, because there is no browser
here. The install is configured as a household's would be: ext4 on /dev/sda (the
32 GB system disk; the 100 GB data disk is never offered), Europe/Budapest,
tester1@felhom.eu, tester1.enkicsifelhom.hu, DHCP values untouched.

Two mechanisms measured rather than assumed, and written down so the next session
does not re-derive them: arrow keys do NOT cycle a value row — Enter opens a list;
and one Up from <Next> lands on a CHECKBOX, so the hostname is five rows up, not one.

The „automatically reboot" box is left ticked on purpose: a volunteer would leave it,
and the trap it causes is already a documented finding.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 17:48:59 +02:00
admin 95e1a39ee3 golden 0.244.0 vouched, floor raised, fresh box installing from the built ISO
gates / gates (push) Successful in 19s
Vouch as a three-field change: golden 0.243.0 -> 0.244.0 with its new checksum;
agent and min_agent stay 0.131.0 because controller 0.244.0 declares the same
MinAgent. Floor raised 0.242.0 -> 0.244.0 with min_agent 0.131.0 so the hub does not
hold it — and it delivered: demo-felhom moved to 0.244.0 by itself within minutes.

VM 335 created from the BUILT 1.28.0 image, disks on /mnt/hdd_1, boot order set in
its own call. The boot menu proves the gate's menu criterion visually: two Hungarian
interactive entries and a 15 s countdown, no Proxmox entry, no automated entry.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 17:35:51 +02:00
admin f57aed9ab9 golden 0.244.0 baked and PUBLISHED — after three attempts, all three failures mine
gates / gates (push) Successful in 23s
sha256 18328a3c7579628b8a7e9639777043db48063c86d37a2e0c221a6ccba6d755a0, 653 609 190
bytes, registry serves it. Markers: overlay2, both mount points, upload OK, no FATAL,
no publish-SKIPPED. Token-leak control passed with a planted positive control.

The three failures are written up because each is a rule this project already has:
scp -p instead of -P (nothing copied); the publisher run without the GITEA_USER it
requires, then the archive destroyed BEFORE checking the outcome; and a rewrite that
dropped the chmod, where the unit reported Result=success while the script inside it
had died on Permission denied.

The fix that matters is the gate: teardown now happens only when the REGISTRY serves
the package — not on an exit code, not on a log sentence. It held: on the failed
attempts the VM and its archive were left in place.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 17:32:18 +02:00
admin 87cb923390 ruling 3 recorded: the restart brake is blind to a slow crash loop (R-539)
gates / gates (push) Successful in 24s
It sits beside the 3-in-15-minutes budget in the host-agent design, because that
sentence and its measured blind spot belong together: four kills 20 minutes apart
were all restarted, none accumulated, and the only trace was an info event that
mails nobody. The budget itself is unchanged.

Also recorded: why Part C.2's re-issue button is correctly hidden for a customer
with no host, the venue's storage reconciliation (nvme-scratch IS /mnt/hdd_1), the
built ISO landing on demo-hp byte-identical, and my own scp/-P mistake that cost a
golden bake — including why its token-leak check reported a false hit on an empty
needle.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 17:17:35 +02:00
admin f2250ca31f register: R-537/R-538/R-536 closed, R-534's grant recorded, three rows opened
gates / gates (push) Successful in 20s
Closed with live proof on demo-hp (controller 0.244.0): the per-tier label and the
restore refusal. R-536 closed with its red-proofs and the hub's two new event types.

R-534 carries the measurement that matters: DatastorePowerUser is Backup+Prune only,
PBS has no custom roles, so DatastoreAdmin at the datastore root for the hub's user
is the narrowest grant that works. The row stays open until a re-issue is seen to
succeed end to end.

Opened: R-539 (a second, slower restart counter — the operator's ruling, for the
nightly), R-540 (one pool box, no selection rule when it fills), R-541 (no path to
move a customer between off-site boxes).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 17:13:51 +02:00
admin ee704b2cf2 evidence: the grant is measured, the off-site default is live, the refusal fires
gates / gates (push) Successful in 20s
ep0: DatastorePowerUser carries Backup+Prune only — measured, not recalled — so the
narrowest role that works is DatastoreAdmin, applied for the hub's user at the
datastore root only. Datastore.Modify now present; the per-customer DatastoreBackup
entries are untouched.

Tester 1's off-site tier is provisioned (shared, 100 GB) and a quota edit REUSES the
same sub-account (311327 all three times), 100 -> 150 -> 100 read back from the form.

R-537 and R-538 proven live on demo-hp running 0.244.0: Paperless's tier-1 row reads
„DB + Konfig" with the new sentence while tier 2 still reads „DB + Konfig + Adatok",
and pressing restore returns the Hungarian refusal with the app untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 17:10:37 +02:00
admin c033b3b617 ISO 1.28.0 source: the console stops showing the pairing code once bound (R-535)
gates / gates (push) Successful in 20s
Measured 2026-09-16: 25 minutes after a successful bind AND claim the console still
showed the pairing code under a line promising the screen refreshes itself.

print_bound_banner is printed the moment the bind delivery lands. It does NOT name
the dashboard URL: the one-shot delivery carries the customer id, passphrase and
mode, not the domain, so naming an address would mean inventing one. The residue —
the console still does not reflect the later CLAIM, because this unit has exited by
then — is recorded in the changelog rather than implied away.

Not published: the built image needs the release gate and the operator's yes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 17:05:35 +02:00
admin 638535b49b hub: deploy v0.116.0 (off-site on by default; the R-536 event pair)
gates / gates (push) Successful in 21s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 16:58:36 +02:00
admin 3738dfc548 hub v0.116.0: every new customer starts WITH the off-site copy (operator ruling)
gates / gates (push) Successful in 19s
Off-site is ON by default for a new customer — shared, 100 GB soft quota prefilled,
the checkbox kept so an operator can opt a customer out. The reason is this repo's
own [FACT]: the whole-guest tiers do not carry the data drive and a Tier-1 unit has
no file leg, so with this unticked a one-drive box keeps NO copy of the household's
own files. Measured on a fresh box the same day.

The quota is prefilled because the fill warning only fires when quota_gb > 0.

Also registers controller v0.244.0's app_deploy_started / app_deploy_failed in both
allowedEventTypes and customerMessages, per the rule that the two move together.

Red-proofed: dropping the default fails the new render test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 16:56:01 +02:00
admin dfd854474e drill 0.243.0 complete: 0 interventions, the connect e-mail proven, NOT ready for a volunteer
gates / gates (push) Successful in 21s
The automatic connect e-mail is proven with a real mailbox: the host record was
deleted at 12:22:59Z and the mail reached the customer at 12:23:00Z, one second
later, with selfbind_link_sent (host delete) on the timeline. The requirement was
two minutes. The hub refuses to delete an ONLINE host with no override, so the
record had to fall stale first — that wait is part of the proof.

Interventions: 0. Every P1 fix this drill set out to prove held on a fresh box.
The verdict is still no, for a new reason: a one-drive box with no off-site tier
keeps none of the household's own files in any backup, the page says otherwise,
and the restore that should save them makes it worse (R-537, R-538).

Teardown, three layers, stated. Customer tester-1 kept; RESET never used; nothing
on the off-site server written or removed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 14:27:04 +02:00
admin ba33db3108 drill 0.243.0: teardown section and the F9'' row in the fault table
gates / gates (push) Successful in 21s
Machine and host layers are done and stated. The hub layer is deliberately waiting:
a host delete is refused while the host is ONLINE, with no override by design, so
the record must fall stale first — that wait is part of the proof.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 14:00:51 +02:00
admin 0b63574293 drill 0.243.0: F9'' answers the supervisor budget question; machine torn down
gates / gates (push) Successful in 20s
Three more kills 20 minutes apart: recovered in 61 s / 41 s / 61 s, and none of
them accumulated, because the window is 15 minutes. Four restarts, zero pauses.
So the brake catches a FAST crash loop and is blind to a SLOW one — a controller
dying every 20 minutes is restarted forever, and the only trace is an info event
that mails nobody. Measured, not changed: the options are written into R-531 for
the operator to rule on.

Machine layer torn down: VM 334 purged with its disks, demo-hp's own containers
9201 and 9202 untouched. Evidence copied off the box first, token-leak control 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 13:59:27 +02:00
admin db58af80a2 drill 0.243.0: interventions counted (0, O1/O2 apart) and the alarm truth table
gates / gates (push) Successful in 21s
Nothing on the walk needed a shell or an operator. The four moments that could be
mistaken for help are listed with the reason each is not one — two of them were my
own errors driving the API, and one was my own damage during the memory test.

The alarm table is now measured from two independent sides: the hub's own log lines
and the inbox. The one-hour operator cooldown is proven to suppress AND to release
(backup_tier_skipped mailed 12:08, suppressed 12:37 and 12:58, mailed again 13:18).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 13:21:17 +02:00
admin f96f93081d drill 0.243.0: Phase 3 morning-after, off-site not walked (R-534), R-528 and R-511 updated
gates / gates (push) Successful in 20s
The off-site restore onto 9202 cannot be walked: this box never had an off-site
tier, because the re-issue fails on the endpoint token's missing Datastore.Modify
grant. Read-only listing of ep0 shows ns/tester-1/ct empty both before and after
the drill, with ns/demo-hp/ct as the positive control. Nothing on ep0 was written,
removed or pruned.

R-528 re-measured on a second, different box: all three OOM signals silent again.
R-511 records that its shipped fix is sound and inert until the grant is given.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 13:18:50 +02:00
admin 725a81a66a drill 0.243.0: Phase 2 faults F10/F11/F12 measured, R-537 and R-538 filed
gates / gates (push) Successful in 20s
F10 (a child deletes the photo folder) is the finding: on a one-drive box with no
off-site tier the household's own files are in NO backup — the whole-guest tiers
exclude mp8 by design and the app's file leg lives at tier 2/3. The app page still
labels tier 1 „DB + Konfig + Adatok" (R-537), and the restore reports success while
leaving Nextcloud listing five photos it cannot open, after wiping the app's own
trash which still held every byte (R-538).

F11: the claim page locks out after the SECOND wrong code (15 minutes), the alarm
fires and is true; Nextcloud does not lock out. F12: two reboots 60 s apart, all
seven stacks back in 124 s, and the supervisor did not count the boots.

Evidence: audits/evidence-drill-0243-2026-09-16/phase2-f10.txt, -f11.txt, -f12.txt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 13:15:41 +02:00
admin bca45aaed1 drill phase 1 complete: fresh box on golden 0.243.0 (sha proven), tunnel, file manager, apps, vaultwarden 400, paperless 20/20, 26 s downtime, absent-tier skip live
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 12:30:47 +02:00
admin 31858d3c52 drill: R-535 filed (console still says 'waiting to pair' after bind+claim); phase 1 evidence — apps, vaultwarden 400, paperless 20/20, per-tier backup page
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 12:27:30 +02:00
admin d1005427a3 drill: fresh box walked — tunnel answers (R-510 closed), file manager generated password, per-tier backup page; adopt blocked by ep0 grant (R-534 filed)
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 12:23:48 +02:00
admin 7eedaac33e drill phase 0: golden 0.243.0 baked+vouched, N100 signed to agent 0.131.0, rulings 1 and 2 recorded; R-529/R-533 closed, R-530 narrowed
gates / gates (push) Successful in 27s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 11:14:10 +02:00
admin 2dd80a5d28 deploy: hub 0.115.0
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 10:54:19 +02:00
admin 926723749d hub v0.115.0: host_* mails skip the quiet hour (ruling 2, R-529); ruling 1 recorded (CC may sign agent_update until the first paying customer)
gates / gates (push) Successful in 23s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 10:53:22 +02:00