8e4365a4c7f11821ed3b12866c8ca53b225932e9
1354 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8e4365a4c7 |
chaos night: the household loop summary for the whole night
gates / gates (push) Successful in 23s
192 probes from 21:06:30Z to 00:18:10Z. Seven lines flagged as failures, of which only TWO are real events - and both are outage pairs caused by an injected accident: wiki during round 2's power cut, cloud during round 10's hard reset. Each lasted less than one probe interval and healed by itself. Three of the seven were my own classifier counting a 301 redirect as a dashboard failure. The log carries the correction in its own words at 21:11:25Z, and the wrong lines were left in place so the correction is visible. Stated rather than glossed: ten of the twelve rounds left no mark in this log at all, including the twenty minutes with the drive pulled. The loop samples each name every two minutes, so that silence is a limit of the instrument, not proof the household saw nothing. Ninth instrument slip recorded: the first listing printed nothing and said 'binary file matches' while the COUNT had already printed, so the summary looked complete while the detail was dropped. Not corruption - zero null bytes, one line of padding spaces. Re-read with grep -a. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
d335033a6d |
chaos night round 12: the closing control round, and the truth table complete
gates / gates (push) Successful in 22s
Steady at t+1s, every door 200 at both readings, household 2 lines 0 failures, and no new alarm - the newest feed entry is still round 11's recovery. Nothing was raised about a box nothing was done to. Checked rather than assumed: inject.sh's default branch exits 2 on an unknown accident, so the control rounds never reach it - the runner handles the no-accident case itself. A broken injector produces exactly the same result as a control round, and only the code path distinguishes them. Alarm truth table, all twelve rounds: 17 alarms fired, 17 true, 0 missing. Three design gaps filed (R-547, R-549, R-550). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
d91822c689 |
chaos night: ep0 baseline and Phase 2 readiness, taken without touching the box
gates / gates (push) Successful in 20s
ep0 baseline (read-only, nothing removed, no prune): three namespaces, each with ct/9201 holding 2 snapshots, six in total, 16G used of 98G. The teardown repeats this listing so 'the backups stayed' is a comparison, not an assertion. Phase 2 readiness: scratch guest 9202 is running with a healthy controller and restic 0.14.0 inside the controller container, but has NO off-site target configured - so the restore will need the box's own repository address and password handed to it. Recorded decision: I did NOT read those from the box while round 12 ran. Round 12 is the closing control round and its whole value is that nothing was done to the box during it. The read costs nothing to defer; the round cannot be re-run. Also recorded: the eighth instrument slip - perl locale warnings plus a head that cut the output before the marker returned an empty block that looked like '9202 has no containers'. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
70bffb1676 |
chaos night round 11: the drive pulled for 20 minutes, and eight true alarms
gates / gates (push) Successful in 25s
The richest alarm round of the night, and every alarm was true and correctly paired: storage_disconnected naming the drive by the household's own label, four app_start_failed naming exactly the four apps whose data lives on that drive, health_degraded, then storage_reconnected and health_recovered. The other eleven apps kept serving throughout. The box recovered unaided in 67 s after the drive was plugged back in, with the front door following at 128 s. The drive came back clean: 98 G, 2% used, mountpoint config unchanged. The round's own snapshot showed health_degraded with no recovery, which would have been the first missing alarm of the night. The recovery had fired seconds after the snapshot. A re-read taken after the precondition found it. Not a missing alarm - a premature reading, caught by the discipline the earlier mistimed readings forced. Truth table now: 17 alarms fired, 17 true, 0 missing, across eleven rounds. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
73ac9d7871 |
chaos night: pre-round-11 steadiness check, and a seventh instrument slip
gates / gates (push) Successful in 23s
The box is steady and the fences are clean: 26 containers, both safety units active and enabled, 0 guard kills, 7556 MB free, all real front doors serving, demo-hp back at -P ACCEPT with 0 physdev rules. The slip: my check probed 'docs', a hostname that does not exist on this box. Traefik's 404 for an unknown Host is correct behaviour, not damage. The real names came from the household loop's own list. A probe with the wrong target produces a confident number that means nothing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
3f844b723a |
chaos night: interventions ledger, built from the evidence not from memory
gates / gates (push) Successful in 19s
One intervention in ten measured rounds: round 6's killed local backup leg. Both pre-declared presses went unused - the automatic self-bind mail was waiting, and the acknowledged-delete path re-issued PBS credentials by itself. Phase 0's seeding repairs are listed separately and in full: that damage was mine, the product behaved correctly throughout, and every repair went through the product's own endpoints. Acts on my own instruments are excluded, with the reason stated, so the count cannot be gamed in either direction. Standing against the stop rule: 1 of 4. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
51782a40b4 |
chaos night: alarm truth table extended to rounds 1-10
gates / gates (push) Successful in 23s
Ten rounds, 9 alarms fired, 9 true, 0 missing. Three design gaps (R-547, R-549, R-550). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
9f40dc3289 |
chaos night round 10: a restore leaves no record, and four of my instruments failed
gates / gates (push) Successful in 21s
The box passed the roughest pair drawn. A hard reset four seconds into a
restore: 26/26 containers back in 150 s, boot reconciliation naming the app it
recovered, every front door serving, one true controller_started alarm, no
false one, no intervention.
R-550 filed (P2): there is no restore record anywhere. Four candidate status
endpoints 404, no restore field in the status JSON, only a button label on the
pages, and no file at all modified in the reset window. An interrupted restore
and one that never happened look identical to the customer. Honest limit
recorded: only four seconds elapsed and the pre-reset log is unrecoverable, so
the absence of a record is what is filed, not a claim about how far it got.
Four instrument faults, all mine, all in the evidence:
* a 'nothing was logged' claim that was unfalsifiable when written - the log
stream holds zero lines before a reset;
* an on-disk check against /opt/felhom/data, a directory that does not exist;
* a household count reporting 0 lines and 0 failures when the truth was one
line and it WAS a failure - the runner now prints both operands;
* the disk guard was a TRANSIENT unit reporting 'active' all night, and was
absent from the reset onward. It is now file-backed and enabled, and its
script is copied off the box for the first time.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
|
||
|
|
f973fd7151 |
chaos night: the hub link repaired itself on the next cycle, and a late ghost task
gates / gates (push) Successful in 21s
Round 9's recovery confirmed end to end: 23:23:43Z 'Hub report pushed successfully (15354 bytes)', exactly 15 minutes after the cycle that failed, with nothing done to the box. The reading was deliberately taken after the report was due so it could not be premature. Also recorded: round 1's poll loop reported 'completed, exit 0' two hours after it stopped working. It went silent during round 2's power cut and its SSH hung until TCP reset it. Three familiar classes in one: silence is not completion, the exit code described the local shell not the remote work, and a very late completion notice can be mistaken for a fresh result. It touched nothing after 21:24:56Z, so no round is contaminated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
70f1e01736 |
chaos night: alarm truth table extended to rounds 1-9
gates / gates (push) Successful in 22s
Nine rounds, 8 alarms fired, 8 true, 0 missing. Two design gaps (R-547, R-549). Also records what this night will probably NOT answer: the event-drop path has never been exercised, because no event was raised during any of the three internet cuts, and rounds 10-12 draw no further cut. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
889310ec17 |
chaos night round 9: what a lost hub report actually costs, measured
gates / gates (push) Successful in 21s
With the injector corrected, the hub really was unreachable. The controller built its 23:08:42Z report, retried the push three times over 1m40.8s and gave up at 23:10:23Z - 31 seconds before the link returned. Nothing queued, which is correct: a report is a snapshot, not a fact. The box passed. 26 containers throughout, every front door serving, both the hub link and the host-agent link repaired unaided the moment the block lifted, no alarm fired and none should have. R-549 filed (P2): the staleness threshold (30 min) is exactly twice the report cadence (15 min), so ONE failed push spends the entire budget. The measured gap was 29m59s - one second inside the alarm. A healthy, self-repaired box came that close to paging the operator. Also recorded: the injected cut is broader than its name - it severed the controller from its own host agent too, which a real ISP outage would not do. The caveat travels with rounds 7, 8 and 9. The event-drop path remains unmeasured, because no event was raised during any cut. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
418f3a2c20 |
chaos night round 8: the accident did not do what its name said
gates / gates (push) Successful in 20s
Round 8 (backup-app nextcloud + internet cut) passed on the product side:
26 containers throughout, LAN doors served the whole ten minutes, public
path restored unaided in <=43 s, no alarm fired and none should have.
The finding is against my own instrument. The hub report due at 22:38:43Z
fell inside the cut and SUCCEEDED, because hub.felhom.eu resolves to a LAN
address (192.168.0.192, measured from guest and host) and the injector
allowed the whole LAN. So rounds 7 and 8 never tested hub unreachability,
and the dropped-event behaviour is still unmeasured.
Two fixes, both to the harness, neither to the product:
* inject.sh now blocks the hub address from the VM's side (the host tap
rule). The hub itself is untouched - the fence is kept. It refuses to
inject at all if it cannot resolve the hub.
* run_round.sh no longer reads the front doors 3 s after an unblock. Every
door reading now states its own timestamp and a second reading is taken
60 s later. This was the sixth mistimed reading of the night.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
|
||
|
|
e45fb5e37f |
chaos night: draft alarm truth table for rounds 1-7
gates / gates (push) Successful in 21s
Built from the verdicts recorded in each round's write-up. 8 alarms fired, 8 true, 0 missing. The one design gap (a transient full disk is never mentioned to anyone) is already filed as R-547. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
b917879e15 |
CHAOS NIGHT round 7 closed: reporting resumed on time, nothing lost
gates / gates (push) Successful in 20s
The post-block hub report fired at 22:23:43Z - "Hub report pushed successfully (10538 bytes)" - and the cadence held to the second across the whole accident: 21:53:49Z, 22:08:43Z, 22:23:43Z, fifteen minutes apart, with a ten-minute network cut sitting between the second and third. So round 7's complete answer: the box lost its way out for ten minutes, kept every app serving at home, restored the public path unaided in ~64 seconds, raised no alarm (correctly - staleness is 30 minutes), attempted no hub contact during the outage because none was due, and then reported on schedule. Nothing was dropped because nothing was sent. Good, and deliberately narrow: this did NOT test what happens to an alarm raised WHILE the hub is unreachable. Rounds 8 and 9 are also ten-minute cuts and one should contain a scheduled report naturally. Also recorded: the fifth mistimed reading of the night, caught this time by the measurement labelling itself - the command printed its own timestamp next to the due time, so a reading taken fifteen seconds early announced itself instead of becoming "the box never resumed reporting". Same principle as the marker blocks that caught the password-file check and the command lines that caught the vzdump self-match: make the instrument say what it actually did. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
3129d4f6b9 |
CHAOS NIGHT round 7: the cut is invisible at home, ten minutes long from away
gates / gates (push) Successful in 21s
Round 7 (use nextcloud, internet cut ten minutes). What the customer saw depends entirely on where they stood: at home nothing at all - traefik answered 301 throughout and every app kept serving; away from home, ten minutes of 502, then 200 again about a minute after the block lifted. The box kept all 26 containers running, needed no repair, and re-established the way in unaided in ~64 seconds. No alarm fired, and none should have: the ladder puts node_stale at 30 minutes and this was ten. The question the round was meant to answer is recorded as NOT EXERCISED rather than passed. Are alarms raised while the hub is unreachable retried and then silently dropped? The box reports every 15m0s (measured: 21:53:49Z, 22:08:43Z) and the cut fell entirely between two reports, so nothing was attempted and nothing could be lost. Rounds 8 and 9 are also ten-minute cuts and one should contain a scheduled report naturally - they are not re-timed to force it. The fence held, checked against a baseline taken BEFORE the round: demo-hp is back to -P FORWARD ACCEPT, zero physdev rules, sysctl 0. That host also carries guests 9201 and 9202, so an abandoned rule would have been a fence breach rather than an untidy drill. Also labelled honestly: the runner's own 530 reading was taken three seconds after unblocking and measures nothing about recovery. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
e61aac1d8f |
CHAOS NIGHT: two enumerated gaps become rows in the same session
gates / gates (push) Successful in 22s
R-547 (P3): a disk that fills and empties between sweeps is never mentioned to anyone. The guest's root filesystem sat at 96% for ten minutes and no alarm of any kind fired - checked twice, once by the round's runner and once independently after the fill was released. disk_critical is defined at >=95% used, but the fill-watch is a DAILY sweep plus one check ~90s after a controller start, so a ten-minute window contains no check. The timing was almost comic: the controller restarted at 21:28 after the previous round's power cut, so its single opportunistic check ran about twenty seconds before the disk filled. This is the ladder working as designed, not a missed alarm - it is filed because the honest answer to "would the household be told?" is no, and that is written down nowhere. R-548 (P3): the whole-guest backup's LOCAL tier cannot fit on a small-system-disk box and retries on that tier for ever. A ~29GB source into a 14GB pve-root, measured falling at ~16MB/s - under four minutes to a full / on the nested PVE. The product's behaviour is correct throughout: it failed the tier, named it, scheduled a retry, its status surface agreed, and the off-site tier then succeeded from the same snapshot in ~8.5 minutes taking no local disk. What is filed is the loop: on a box this shape the local tier can never succeed. Honest caveat recorded in the row - the 32GB system disk is this drill's own fixture choice - but nothing checks whether the local target could hold the source before starting. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
3abd25e681 |
CHAOS NIGHT round 7: a ten-minute outage falls between two reports
gates / gates (push) Successful in 21s
Measured from the box's own log rather than recalled from documentation: 21:53:41Z Registered periodic job: hub-report (every 15m0s) 21:53:49Z Hub report pushed successfully (30535 bytes) 22:08:43Z Hub report pushed successfully (15934 bytes) - exactly 15m later The block runs ~22:10Z to ~22:20Z and the next report is due ~22:23:43Z, after it lifts. So the box never attempts a push while cut off: nothing was tried, nothing failed, nothing was lost. The honest verdict for the question I wanted this round to answer - are alarms raised while the hub is unreachable retried and then silently dropped? - is NOT EXERCISED, not "passed". Recorded that way. It is still a finding of its own: a ten-minute internet outage is invisible to the fleet view because the box had nothing due to say, and the hub's staleness threshold (30 minutes) is set well beyond it. The two mechanisms agree. Rounds 8 and 9 are also ten-minute cuts at ~25-minute spacing against a 15-minute cycle, so one will very likely contain a scheduled report and exercise the drop behaviour properly. They are NOT re-timed to make that happen - re-timing a round to get a better result is choosing the night after the fact. Also captured while blocked: internet unreachable from the box, LAN reachable, 26 containers up, free space unmoved, disk guard silent. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
a103b62330 |
CHAOS NIGHT round 7: the block works, and "vzdump procs: 2" was my own command
gates / gates (push) Successful in 20s
Mid-cut measurements, asked of the box while its network was blocked: from the box: internet blocked (ping 1.1.1.1 fails), LAN reachable 26 containers still running - the apps do not care the internet is gone free / unchanged at 7573M; diskguard active with zero log lines from DooPlex: LAN 301, public 502 The two paths separate cleanly, and the public failure code differs from round 4's on purpose: 530 when cloudflared was dead (Cloudflare had no tunnel at all), 502 now (the tunnel lives but can reach nothing). Two different failures of the same journey, reported differently without being asked to. And the twelfth self-inflicted reading of the night, corrected: every "vzdump procs: 2" was MY OWN COMMAND. The [v]zdump bracket trick stops the pattern matching itself, but the label I echoed - "vzdump procs:" - contains the word, so ps listed my own shell and the grep counted it. No second backup was ever running; the box has been idle since the off-site leg finished at 22:08:27Z. Same family as the pkill -f that killed my own watcher earlier: a pattern that matches the hand holding it. The cure that worked: ask for command lines, not a count. The disk guard stays regardless - the local tier really did announce a retry with backoff, that retry is still due, and the guard has cost nothing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
c3722e06d2 |
CHAOS NIGHT round 6: the local backup tier cannot fit, and the box says so properly
gates / gates (push) Successful in 21s
Round 6 (whole-guest backup, control round). The local tier failed - a ~29GB source into a 14GB root filesystem - and the product handled it well: whole_guest_backup_failed (error): "Whole-guest backup FAILED on the LOCAL TIER - retrying with backoff (next attempt in 15m0s)" It names the tier rather than "the backup", says what it will do next, and its status surface agrees (target_id local, success false, size_bytes 0). Then the OFF-SITE tier ran from the same snapshot with the apps already back up, finishing in ~8.5 minutes, encrypted to ep0, consuming no local disk at all. During the backup 4 of 26 containers were up and every app answered 404 publicly - and NO alarm fired for those stops, which is correct: the backup's own stack stops are suppressed, so the box does not alarm about downtime it caused deliberately. The downtime number (~5m43s) is recorded as CONTAMINATED by my own intervention rather than presented as clean: I killed the local leg partway through. A safety guard now runs on the box before round 7, declared and not a measurement: the failed local tier retries every ~15 minutes and would consume / at ~16MB/s while round 7 has the network cut. The guard kills only a local-tier dump, only below a 2500M floor, and logs every action. If it fires it is an intervention and will be counted as one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
aaf0537665 |
CHAOS NIGHT round 6: the off-site leg takes no local space - measured, not argued
gates / gates (push) Successful in 20s
While the off-site leg streams to ep0, free space on / does not move: 22:02:09Z procs=2 apps=26 free=7573M 22:02:40Z procs=2 apps=26 free=7573M 22:03:12Z procs=2 apps=26 free=7573M That is the difference between the two legs of the whole-guest backup, stated as a measurement. The local leg consumed ~16MB/s of / and would have filled it in under four minutes; the off-site leg has run for several minutes and taken nothing, because it streams encrypted to ep0 rather than writing an archive locally. All 26 apps are back up and serving while it runs. So the whole-guest backup is not "too big to work" - it is too big for the LOCAL target, and the tier that matters for disaster recovery is unaffected. My first conclusion was wider than the evidence; this narrows it. ep0 is deliberately not queried to watch the snapshot land: the box's own task status answers the same question when the leg ends, and tonight's fence is that ep0 is written only by the product's own path and read only when nothing else can answer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
36ae3b3191 |
CHAOS NIGHT round 6: I stopped a LEG, and the box carried on by itself
gates / gates (push) Successful in 21s
Correction to my own account, recorded where the wrong version stood. I wrote "I stopped the backup". What I stopped was its LOCAL leg: the task list shows that job ending 23:59:45 with status "job errors" - my pkill - while a second vzdump was already running, streaming encrypted to ep0 (--repository felhom@pbs!tester-1@...:felhom-offsite --ns tester-1). The arithmetic that forced the intervention still stands: the local leg was writing a ~29GB source into a filesystem with 3.6GB free, falling at ~16MB/s, which gave under four minutes before / filled and the nested PVE wedged - the Phase 0 failure one level up. But the off-site leg needs no local space at all, so the box's design copes with exactly the problem I thought I was rescuing it from. It is still counted as an intervention: I reached in and killed a job. The box then recovered unaided: 22 containers at 22:00:52Z, 26 at 22:01:13Z, and / went from 2539MB free back to 7573MB once the partial archive was removed. The older completed backup was left untouched. Round 6's valuable half stands: during the backup 4 of 26 containers were up, every app returned 404 through the public route, and NO alarm fired - the suppression of the backup's own stack stops held. And the waiting rule for the off-site leg is fixed in advance: round 7 does not start while it runs, unless it is still running at 00:45, in which case round 7 proceeds and records that it cut an in-flight backup - labelled as that, not as a clean internet-cut round. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
d4318529b4 |
CHAOS NIGHT round 6: what a whole-system backup costs the household
gates / gates (push) Successful in 20s
Measured while the vzdump was in flight (started 21:55:24Z, 2.1GB and growing): - containers running: 4 of 26. The whole-guest backup stops twenty-two apps. - front doors: LAN 301 (traefik is one of the four still up) but public 404 - nothing behind the proxy to serve. Every app unavailable for the duration. - alarms: NONE. Twenty-two apps went down at once and not one alarm fired. Those two findings point opposite ways and both matter. The downtime is real and total, not a brief pause - this is the standing whole-system-backup downtime row, seen on a fresh box with twelve apps. And the suppression is correct: the backup's own stack stops belong to a suppression set, so the box does not alarm about downtime it caused deliberately. The duration itself is taken from the round's runner when it reports, not estimated from a single sample. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
eb638d303b |
CHAOS NIGHT rounds 4-5: the box repairs its own tunnel, and survives docker dying
gates / gates (push) Successful in 22s
Round 4 (tunnel killed 10 min): cloudflared came back in ~97 seconds and the CONTROLLER did it, not Docker - RestartCount=0 proves the unless-stopped policy never acted, and the log shows the protected-infra recovery redeploying it. health_critical fired at 21:43 and health_recovered closed it at 21:48, so the alarm was not a dead end. Round 4's drawn ACTION never ran, and that is recorded rather than re-run: the round-2 power cut rebooted the guest, /tmp is cleared on boot, and the dashboard password lived there. The off-site run was never triggered. Re-running a round after watching it fail is how a drill starts choosing its own results. The file now lives in /root, so round 10's hard reset cannot disarm rounds 6, 8 and 10. Round 5 (docker restarted): all 26 containers back in 16 seconds, doors serving again in ~40, controller_started true and correct, nothing missed. The household loop is recorded as NOT SAMPLED - 0 lines because the round was shorter than its 2-minute sampling interval, which is not the same as 0 failures. Two traps avoided and written down: round 5's own alarm snapshot was taken six seconds after the controller started, from which controller_started looked missing (it fired); and my check for the password file put its redirect on the wrong host, reporting "not there" for a file that was present all along. Tenth and eleventh of the same family tonight - a check whose own precondition was wrong, answering confidently. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
3e6645413e |
CHAOS NIGHT: round 3 written up, and the household loop's blind spot stated
gates / gates (push) Successful in 21s
Round 3 in the findings document: the disk sat at 96% for ten minutes, twelve apps kept serving, the household loop logged 12 operations with zero failures, and NOTHING was ever raised. The silence is the finding, and it was predicted from the ladder before the round: the fill-watch is a daily sweep plus one check ~90s after a controller start, and that single check ran about twenty seconds before the disk filled. Recorded with it: I twice labelled a mid-window reading "end of window", estimating the clock instead of reading it. The readings were unchanged but the label was wrong, and "nothing yet" is not "nothing ever". And a limit of my own instrument, stated before its numbers get quoted: the household loop does not follow redirects, so it measures "is the app serving on the box" and never "can the household reach it from outside". It logged zero failures straight through round 4's tunnel outage while the public route was returning 530. So "0 household failures in round 4" must not be read as "the household was unaffected" - someone away from home would have met 530 for about ninety seconds. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
ec84eadc19 |
CHAOS NIGHT round 4: the box repairs its own tunnel in 97 seconds
gates / gates (push) Successful in 22s
Killed cloudflared and deliberately never restarted it, because whether it returns by itself IS the measurement. It returned in ~97 seconds (killed ~21:41:30Z, running again 21:43:07.478Z), and the box did it, not Docker: RestartCount=0 proves the unless-stopped policy never acted, and the controller log shows "[infra] deploying cloudflared -> /opt/docker/stacks/cloudflared". That is the product's protected-infra recovery repairing one of its own infrastructure stacks unasked. health_critical (error) fired at 21:43 - "Rendszer allapot kritikus (volt: ok)" - which is exactly what the ladder predicts for a missing protected container, and it is true. Whether health_recovered closes the pair is checked at the end of the round, not guessed at now. A METHOD CORRECTION that retro-labels every front-door reading tonight: the 530 during the outage is a Cloudflare status, which exposed that `curl -sL` was following traefik's 301 out to the public hostname and back down the tunnel. So every "front door" reading so far measured the PUBLIC path, not the LAN. It does not invalidate the readings - a 200 by that route proves more, not less - but it invalidates the label, and with it any claim of the form "the app is fine, only the tunnel is down". The two paths are now measured separately: during this outage the public route gave 530 while traefik answered 301 locally throughout. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
34d22a1a92 |
CHAOS NIGHT round 3: the disk fills for ten minutes and nobody is told
gates / gates (push) Successful in 21s
Round 3 (use bookstack, system disk held at 96% for ten minutes): - the household saw nothing wrong: wiki, status and paste all answered 200 before, during and after, and the background loop logged 12 operations with ZERO failures on a 96%-full disk - the box kept all 26 containers running and released the space cleanly (29G used -> 944M used) with the thin pool untouched at 39.69% throughout - NO alarm fired at any point, checked twice independently after the fill was released That silence is the finding, and it was predicted from the ladder before the round rather than discovered after: disk_critical is defined at >=95% used, but the fill-watch is a daily sweep at 03:30 plus one check ~90s after a controller start. The controller happened to restart at 21:28, so its single opportunistic check ran about twenty seconds BEFORE the disk filled. A disk that fills and empties between sweeps is invisible - by design, but the honest answer to "would the household be told?" is no. Also fixed and explained: my injector printed "unexpected EOF" while the accident itself completed. bash -n passes, so it was not local syntax - G() flattens its argument through `pct exec`, so a nested bash -c '...' has its quoting re-parsed remotely. Both instances were in the disk branch only; the accidents still to come use plain commands. The experiment was verified on the box, not from the script's own account. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
aca0172efd |
CHAOS NIGHT: fences clean, firewall baseline taken, and a third mistimed label
gates / gates (push) Successful in 22s
Fence check mid-night: demo-hp carries only this drill's VM 336; guests 9201 and 9202 are running with their full standing sets intact, bentopdf included. Nothing of theirs was stopped, removed or redeployed tonight. Firewall baseline recorded BEFORE the rounds that need it (7-9 block the box's internet by flipping a host sysctl and inserting two physdev rules): iptables -S FORWARD -> "-P FORWARD ACCEPT" and nothing else physdev rules -> 0 net.bridge.bridge-nf-call-iptables = 0 A control taken before the experiment, so that "it looks clean afterwards" can be a measurement rather than an assertion - on a host that also carries the two standing demo guests. And a third mistimed reading of my own, recorded: I labelled a 21:36:00Z check "end of window" when the fill runs to ~21:39:49Z. The readings are unchanged (nothing fired), but the label is the point - "nothing yet, five minutes in" and "nothing in the whole window" are different findings. From here the end-of-window check is taken when the round's runner reports completion, because the runner knows when it released the fill and I was guessing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
ee3da86d33 |
CHAOS NIGHT round 3: interim reading, and a mislabel caught before it stood
gates / gates (push) Successful in 20s
Five minutes into the 96%-full disk: the apps keep serving (26 containers), the pool is untouched at 39.69%, the filesystem is writable, and nothing has been raised - the newest event is still controller_started from 21:28. I nearly filed this as the "end of window" check. The fill began 21:29:49Z, so the ten-minute hold runs to ~21:39:49Z and this reading was taken at 21:34:23Z, halfway through. It is recorded as INTERIM and the end-of-window check stays owed, because an alarm arriving late is a different finding from one that never arrives - and a mid-window reading standing in for the final one would have quietly turned "not yet" into "never". What it already establishes: twelve apps keep running and serving with the system disk at 96% full, and after five minutes nobody has been told anything. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
fa1ddd92a5 |
CHAOS NIGHT: the internet-block accident now cleans up unconditionally
gates / gates (push) Successful in 20s
Reviewed before it runs unattended at ~01:11, not after. The accident flips a host-wide sysctl on demo-hp and inserts two FORWARD rules, and its cleanup ran only on the happy path: if the script were killed during its ten-minute sleep, or the SSH dropped, the rules and the sysctl would have stayed. demo-hp is a Tier-0 host that guests 9201 and 9202 also live on, so an abandoned FORWARD rule is a fence breach rather than a measurement. It now traps EXIT, INT and TERM, removes both rules and restores the sysctl whatever happens, and clears the trap on the normal path so the cleanup does not run twice. The rules still match --physdev-in on this VM's own tap, resolved at run time, and the default FORWARD policy is never touched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
7221367f38 |
CHAOS NIGHT round 3: the disk fills, and nothing is told about it
gates / gates (push) Successful in 21s
Measured while the guest's root filesystem was held at 96%: - the app's view is a genuinely full disk (29G used, 1.5G free, 28G fill file) - the shared thin pool stayed at 39.69% - fallocate reserves blocks without writing them, so this round is NOT a repeat of the pool exhaustion that wedged the box in Phase 0. Recorded explicitly, because "disk 95% full" invites exactly that wrong reading. - the filesystem stayed writable (a real touch, not the mount flags) No disk alarm fired, and that was PREDICTED from the ladder before the round: the fill-watch is a daily sweep at 03:30 plus one check ~90s after a controller start. The timing is sharper still - the controller restarted at 21:28 after round 2's power cut, so its single opportunistic check ran about twenty seconds BEFORE the disk filled. A disk that fills and empties between checks is invisible; that is by design, but it is the honest answer to "would the household be told?" - no. Household lines are attributed to the right round: the two UNREACHABLE entries at 21:27:57Z are round 2's recovery tail, not round 3's accident. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
bca013edec |
CHAOS NIGHT: round 2 passed, and the OOM finding goes on the row that owns it
gates / gates (push) Successful in 23s
Round 2 (restore gokapi + power cut, 20s into the restore): the box came back BY ITSELF in 148 seconds, 0 -> 25 -> 26 containers, and gokapi - the app being restored when the plug came out - returned healthy. The only alarm was controller_started, which is what the ladder expects for a 60-second outage: no node_stale (30 min threshold), no app_start_failed (90s boot grace). No false alarm, none missed. Round 2's household measure is recorded as NOT COLLECTED, not as a pass: the loop died with the box and zero lines is not zero failures. The round also handed over immich's whole diagnosis. app_oom fired - "immich (immich-postgres) - egy folyamatat a memoriakorlat leallitotta" - naming the app and the exact container. That is why immich saw CONNECTION_CLOSED and crash-looped twelve times. It is added as tonight's line on the EXISTING OOM row rather than filed as a new one, because this project's standing finding is that those signals are invisible inside LXC guests and on this box the scan caught one. The diagnosis I spent twenty minutes reaching from logs was sitting in the alarm feed, correctly labelled, the whole time. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
5b6e4b5c30 |
CHAOS NIGHT: round 2 under way, and the household loop made to survive accidents
gates / gates (push) Successful in 21s
Round 2 (restore gokapi + power cut) is running: the restore started, the plug came out 20s later, and the box is recovering on its own. Found at round 2 and fixed for the rest of the night: the background household loop ran as a TRANSIENT unit on the VM, so the first accident that could have produced household failures - a power cut - instead killed the loop and produced no lines at all. Zero lines is not zero failures, and round 2's household measure is recorded as NOT COLLECTED rather than as a pass. It is now a real systemd unit with Restart=always, enabled at boot, so it returns with the box after the power cuts, hard reset and docker restart still to come. Machinery for the remaining rounds written in advance rather than mid-round: one generic runner covering every action and accident the seed actually drew, so no round is measured a different way from another. It records the same five things each time, counts the household loop's lines and failures for its own window, and names immich as a known pre-existing failure so nothing later is misattributed to an accident. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
cc87efa235 |
CHAOS NIGHT: household whole (11/12), injector fixed, round 2 running
gates / gates (push) Successful in 21s
The household the accidents actually hit: 11 of 12 front doors serve 200, verified with a no-such-host control so a 200 means a real route to a real app. bookstack's 500 was my seventh error and the first I confirmed before acting: the repair script DECODED the stored APP_KEY and printed "prefix ok: False, body length: 96, NOT valid base64" - Laravel could never have used it. A clean redeploy with base64:$(openssl rand -base64 32) had it healthy in 45 seconds. immich is diagnosed (CONNECTION_CLOSED to its postgres during reverse-geocoding init, RestartCount=12) and deliberately LEFT BROKEN: no round in the drawn schedule acts on it, and chasing the one app the schedule never touches would cost rounds that were drawn before the night began. It is named as a known pre-existing condition so no later failure is misattributed to an accident. Machinery fixed before its round arrives, not during it: - inject.sh used `qm guest exec`, which this box cannot do (no guest agent, measured in Phase 0). Three drawn accidents depend on in-guest work, so it now goes over SSH + pct exec. - the disk-95%-full accident gained a POOL GUARD: the guest's disks are thin provisioned over the pool that hit 100% and remounted the box read-only earlier tonight, so the fill is capped and any cap is declared in the round's own evidence rather than silently applied. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
a046db7df7 |
CHAOS NIGHT: household verified, and a seventh error of mine found by control
gates / gates (push) Successful in 21s
Verified after the repairs: 26 containers running, all twelve apps deployed, ZERO auth-failure lines on the rebuilt DB-backed apps, and 10 of 12 front doors answering 200 - including cloud (nextcloud) and share (gokapi), both of which were broken an hour ago. The cures are confirmed at the front door, not by a health badge. bookstack diagnosed properly rather than guessed at: its healthcheck exits 22 (curl's "server returned an HTTP error"), a direct request to the container returns 500, its migrations completed cleanly and it has no auth failures. So neither the database nor the image is at fault - the app itself errors. The likely cause is mine: I passed APP_KEY=base64: plus 32 random alphanumerics, which is not a base64-encoded 32-byte key. The repair decodes the stored key and prints its true length BEFORE redeploying, so the hypothesis is confirmed or refuted in the evidence. Also recorded: my sixth slip, running docker over SSH on the VM instead of inside the guest, which printed a tidy table of "absent" and "0" that read like "nothing is wrong" and was produced by a shell with no docker at all. Six of my errors tonight share one shape - a command whose precondition failed, still printing a confident answer - and the same discipline caught every one: ask the box directly, with a control. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
c3e1986aaf |
CHAOS NIGHT: household repaired, headroom checked, round 2 armed
gates / gates (push) Successful in 21s
The five apps my disk-full burst broke are removed with their data and
redeployed cleanly, one at a time. Two distinct faults of mine, with different
cures, and separating them is what made either fixable:
- image layers written while the pool was full -> "invalid ELF header",
exit 127; cured by dropping the image so compose re-pulls
- my re-seed's FRESH database passwords over volumes initialised with the
first set -> Postgres auth_failed / MariaDB "Access denied"; cured only by
removing the app with its data and deploying once
gokapi proves they are different: a new image left it Restarting(1), a new
database made it healthy.
Headroom measured so the disk-full failure cannot quietly repeat: pool 39% of
75.8G, docker filesystem 29%, data drive 1%, 4.3G guest RAM free.
Round 2's script hardened before it runs unattended: its "steady" test now
compares against the count it measured itself in the same round, instead of a
hardcoded 24 that the rebuilt household might never reach.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
|
||
|
|
a1a57ea6d5 |
CHAOS NIGHT: round 1 recorded, and three of my own conclusions corrected
gates / gates (push) Successful in 20s
Round 1 (offsite-run, control round) is written up with its five things: the run worked end to end in 1m45s, and on this REBUILT box the remote repository is orphaned - the documented behaviour, surfaced honestly instead of reported as a successful copy. Three alarms fired, all true and precise. Corrections to my own earlier claims, each recorded where the wrong version was written: - "the 404s were my mistimed sweep" - wrong for four of five. Proven with a negative control (a no-such-host request returns the identical 404, 19 bytes) that traefik simply has no route to an unhealthy container. - "nextcloud is repaired" - wrong. The re-pull fixed the corrupt library, but the app still cannot reach its database, and the container reports HEALTHY the whole time. A health signal is not a data signal. - "all the broken apps are corrupt layers" - wrong. bookstack logged a clean startup, gokapi logged nothing, and immich shows a Postgres auth_failed. The real cause of most of it is mine: my re-seed generated FRESH database passwords over volumes whose databases were initialised with the first set. The affected apps are being removed with their data and redeployed cleanly. Also recorded: an HTTP 000 is "no answer", not "it failed" - the action still took effect. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
5f2ccec64c |
CHAOS NIGHT: household seeded, escrow done, round 1 measured
gates / gates (push) Successful in 19s
Phase 0 finished: all twelve apps deployed one at a time (the parallel burst was my error and is recorded as such), escrow ceremony completed after the box converged the PBS descriptor by itself (~17 min), and the R-543 bar disappeared for good once escrowed. Round 1 (offsite-run, control round, no accident): - the run started on the button's own endpoint and walked every app with real per-volume byte counts (stop -> dump -> restart) - three alarms fired, all TRUE and precise: app_start_failed named the one crash-looping app, backup_run_failures said "1 of 12 ... nextcloud", and offbox_repo_orphaned reported the documented rebuild behaviour rather than claiming a successful copy - nextcloud's crash loop is MY damage (image layers written while the thin pool was 100% full -> "invalid ELF header"), not a product defect, and is recorded that way The catalog bump was prepared and then REVERTED UNPUSHED: the catalog repo's own gates returned INCONCLUSIVE (the volume-persistence prober failed its own canary), and undetermined is never a pass. Round 7's `update` therefore becomes `use`, decided now rather than improvised at 02:00. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
9fae6dfa98 |
CHAOS NIGHT phase 0: golden 0.245.0, a self-installing box, and R-546
gates / gates (push) Successful in 23s
The schedule was drawn from seed 20260917 and written into the findings document BEFORE round 1, with its re-draw log. Phase 0 measured: - golden 0.245.0 baked, published (registry 200, not an exit code) and vouched; the box installed itself from the published ISO 1.28.0 and landed on it with no hand upgrade (controller 0.245.0, agent 0.131.0). - ZERO operator presses: the waiting self-bind mail worked, and the acknowledged -delete path re-issued off-site AND PBS-DR credentials by itself (pbsdr_auto_reissue) - the F-14 half nobody had watched happen live. - R-546 filed (P2): tonight's own guide sends the household to create the recovery code ~17 minutes before the box can do it. It self-heals; the bar urges them there the whole time. Measured on both sides, not inferred. - R-543 proven through its whole lifecycle on a fresh box: bar present while paused, gone for good once escrowed. - Known rows met and recorded, not re-filed: R-542, R-536's failure events. Also recorded honestly: three harness errors of mine (a script that announced "all twelve deploys ACCEPTED" without checking, a "login ok (csrf 0)" that turned eleven of my own 401s into what looked like product refusals, and a head -12 that hid a disk), and a near-miss where I almost filed a defect against a drive gate that was working and logging at DEBUG. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
d124c77e17 |
R-543 closed: the household is asked for the recovery code (controller v0.245.0)
gates / gates (push) Successful in 21s
The tier-3 pause is the zero-knowledge escrow design and is untouched. What was missing was the ASK, while the backup page promised the copy that had never run. - VOLUNTEER-first-hour.md: a new step 6, right after the dashboard password and before the first app - what the code is, where, write it on PAPER, and that Felhom cannot get it back for them. Sections 6..12 renumbered to 7..13. - day0-install.md A.2b: the operator step for a REBUILT box, which was missing. Acknowledged delete -> the hub re-issues by itself; otherwise ONE press of "Re-issue PBS credentials" (F-14 ruling 2026-07-13, hub/internal/web/pbsdr.go). This is the correction to last night's "zero presses" note. - 07-backup-architecture.md: 6.1 records tier-3's paused state as a DESIGN, and 2 records that the household is asked from first login. - capability map: the first-hour row's last gap closed, with what it still does not claim (no volunteer has walked the ask from the written guide). - register: R-543 CLOSED with the live measurements; R-545 filed (nothing un-configures an off-site target). R-511 was already closed yesterday. - STATUS: the answered publish question removed (1.28.0 is live), readiness yes. - evidence: red-proofs, the two-box live validation, teardown on three layers, and both of my own mistakes in this session. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
1acd693854 |
the volunteer's own path verified: felhom.eu/letoltes serves 1.28.0
gates / gates (push) Successful in 21s
The page names 1.28.0 four times and 1.27.1 zero times, the link answers 206, and the published checksum is the checksum of the file that was gate-checked and installed tonight. CI green for both repos, matched by head_sha (669, 648). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
832218dca4 |
ISO 1.28.0 PUBLISHED on the operator's yes; download page and R-535 updated
gates / gates (push) Successful in 20s
Uploaded with env-only credentials and verified by ROUND TRIP: the downloaded bytes checksum to a4cd9b6d…, identical to the built file, and the checksum file is served. 1.27.1 stays in the bucket; nothing was overwritten. The download page now names 1.28.0 with the published checksum (BOM preserved, site gates green). R-535 closes with an honest caveat: the new banner ships byte-identical to repo HEAD and the string is in the published payload, but it was never seen on a screen — the box bound itself while the walk was headless. Also corrected: the 1.27.1 heading still said NOT PUBLISHED although it went out on the big night. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
3f7ac8ee6e |
the backup promise is kept: photos deleted and returned byte-identical
gates / gates (push) Successful in 20s
The capability map's journey row now carries the half it could never finish: five photos in, deleted the way a child would, the old route refusing and touching nothing, the off-site restore returning them, and them opening — sha256 identical, 5 of 5, with a negative control. Stated with it, because both are true: the bind needed ZERO operator presses (the box registered itself and used the mail the hub sent itself), but the PBS cascade needed ONE — the Re-issue press R-511 documents, which then succeeded because of this morning's ep0 grant. R-543 (P1) is the honest caveat: off-site ON by default is not off-site WORKING on day one — a fresh box waits at „Kulcsletétre vár" until the household creates its recovery code, and nothing asks them to, while the tier-1 row already promises that copy. R-544 records a log line that says „escrow deleted" where the effect is demotion to retained custody. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
c18efc0610 |
teardown layer 1, a proper secret-leak check, and the fresh-box proof in both rows
gates / gates (push) Successful in 20s
VM 335 purged with its disks; demo-hp's own containers untouched; evidence pulled off the box before the destroy, with the one thing I could not collect stated (the agent journal — root SSH is refused on the appliance by design). The leak check redone properly: six real secret VALUES as needles against all 41 evidence files, planted positive control matched 6/6, committed evidence matched 0. The earlier „22" was the word „password" in labels — a word count, not a leak check. R-537 and R-538 now carry the fresh-box proof: the labels on a box where off-site is on, the refusal that pointed at the off-site route, and five photos returned byte-identical. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
c91c1377bc |
THE PHOTOS OPEN — the backup promise is true on a box that installed itself today
gates / gates (push) Successful in 22s
Five photos in, deleted the way a child would, and back again: HTTP 200 with bytes 200000/400000/600000/800000/1000000 and sha256 identical to the originals, five of five, with a negative control. Controls at the same moment: status.php 200, WebDAV 207. The two halves that make it honest: - the OLD route refused and touched nothing („a fájlok így a helyükön maradnak"), naming the route that could help; the app was running before and after; - the off-site restore ran in two steps — a verification copy that states „A meglévő adatok változatlanok", then a reconstitution whose message counts „5 fájl és 3 adatkötet és az adatbázis". This morning the same deletion ended with five photos listed, none of them openable, and a success message over the top. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
c5df1a372b |
the off-site tier works on the fresh box, and the photos are deleted for the test
gates / gates (push) Successful in 20s
The escrow ceremony ran (start re-authenticates, phase done, uploaded, sealed) and the one-time recovery code was captured into a 0600 file at the moment it appeared — it is shown once, the box stores it nowhere, and it appears in no committed file. Tier 3 then reported „Sikeres restic → …your-storagebox.de · Helyreállítási egység, titkosítva", and „Kulcsletétre vár" is gone. The restore wizard renders for Nextcloud, which it only does for an app the store can actually restore — checked BEFORE deleting anything, because deleting with an unproven copy is the harm itself. Then the child's action: DELETE /Fotok -> 204, the folder 404s, the photo 404s. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
a73abf04db |
R-534 and R-511 CLOSED — the ep0 grant proven end to end on a fresh box
gates / gates (push) Successful in 20s
The rebuilt-customer case reproduced by itself: the WG-registration hook refused exactly as R-511 describes and named the Re-issue action. Pressing it then worked — reissue ok, pbsdr ADOPTED (gen 2), and the box consumed the single-use secret two seconds later. No permission error. This morning the identical action returned „missing Datastore.Modify … status 255" and a 502. The only change in between is the narrow grant on ep0, and the narrowest role was measured rather than recalled: DatastorePowerUser carries Backup+Prune only, and PBS has no custom roles. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
63e2de9b9e |
R-543: off-site ON by default is not off-site WORKING on day one
gates / gates (push) Successful in 21s
Measured on the fresh box: tier 3 sits at „Kulcsletétre vár" — the off-site copy is paused until the household performs the key-escrow ceremony, and nothing asks them to. POST /backup/offbox/run returns 302 and produces no snapshot; the controller log shows only offsite-credential-retry. That matters more after today, not less: the new default exists because a one-drive box otherwise keeps the household's files in no tier at all, and the tier-1 row now prints „Az alkalmazás fájljait a távoli másolat … védi". On day one that sentence promises a copy that does not exist yet. The good half, proven on the same page: tier 1 reads „DB + Konfig" with the new sentence, and „DB + Konfig + Adatok" appears zero times — R-537 holds here too. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
0cfbfe709c |
R-536 proven live on a fresh box: acceptance no longer claims the app is installed
gates / gates (push) Successful in 20s
At the moment of the 202 the hub received app_deploy_started — Alkalmazás telepítése elindult: Nextcloud — and app_deployed is absent while the install is still running. This morning the same moment produced „Alkalmazás telepítve: Mealie" for an install that was killed five seconds later and never happened. Also recorded: the deploy was first refused 400 because the admin password field is mandatory at the server while its own metadata says required:false with generate:password:16 — the browser fills it with the Generálás button, so a household never meets it, but an API caller that trusts the metadata does. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
2687819195 |
the drive was fine and my instrument was not — corrected, and R-542 filed
gates / gates (push) Successful in 20s
I read „formatted but not mounted" off /api/disks/candidates and briefly held it as a product fault. The storage page — the surface a household opens — says the opposite and is right: Adatlemez, /mnt/felhom-drives/adatlemez, default, active, ext4. R-542 records the real (small) defect: that endpoint offers a REGISTERED, in-use drive under „initialize", with already_mounted null. The page filters it out, so no customer sees it; it fed a formatting flow and it misled a session, which is enough. Also recorded: off-site is LIVE on this fresh box by default — „Aktív — nincs kijelölt alkalmazás" — an hour after that default shipped, with nobody pressing anything. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS |
||
|
|
4209908415 |
fresh box claimed, and it landed on the golden this task published
gates / gates (push) Successful in 21s
Login with the new password returns 302 and the dashboard opens: the claim took.
The page reads controller 0.244.0 — so a box installed from the built ISO lands on
the vouched set with no hand upgrade (agent 0.131.0, controller 0.244.0, PBS wrapper
matching).
Recorded alongside: a transport fact (the same claim page 403s to a python client and
200s to curl seconds apart — the edge judging the client, not the box refusing), and
that the drive init's {"started":true} is an attempt, not a result, so nothing is
deployed onto the drive until the mount is observed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
|