CHAOS NIGHT phase 0: golden 0.245.0, a self-installing box, and R-546
gates / gates (push) Successful in 23s

The schedule was drawn from seed 20260917 and written into the findings document
BEFORE round 1, with its re-draw log.

Phase 0 measured:
- golden 0.245.0 baked, published (registry 200, not an exit code) and vouched;
  the box installed itself from the published ISO 1.28.0 and landed on it with
  no hand upgrade (controller 0.245.0, agent 0.131.0).
- ZERO operator presses: the waiting self-bind mail worked, and the acknowledged
  -delete path re-issued off-site AND PBS-DR credentials by itself
  (pbsdr_auto_reissue) - the F-14 half nobody had watched happen live.
- R-546 filed (P2): tonight's own guide sends the household to create the
  recovery code ~17 minutes before the box can do it. It self-heals; the bar
  urges them there the whole time. Measured on both sides, not inferred.
- R-543 proven through its whole lifecycle on a fresh box: bar present while
  paused, gone for good once escrowed.
- Known rows met and recorded, not re-filed: R-542, R-536's failure events.

Also recorded honestly: three harness errors of mine (a script that announced
"all twelve deploys ACCEPTED" without checking, a "login ok (csrf 0)" that
turned eleven of my own 401s into what looked like product refusals, and a
head -12 that hid a disk), and a near-miss where I almost filed a defect
against a drive gate that was working and logging at DEBUG.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-16 22:56:08 +02:00
parent d124c77e17
commit 9fae6dfa98
46 changed files with 2746 additions and 0 deletions
@@ -0,0 +1,159 @@
# DRILL — CHAOS NIGHT: random actions on random apps while random things go wrong (2026-09-16/17)
**Interventions: PENDING — the run is in progress.**
**Ready for a volunteer: PENDING.**
**The accident-plus-action pair that hurt most: PENDING.**
> **Baselines, verified live against Gitea at 21:49 CEST 2026-09-16 (not copied from the brief):**
> felhom-controller `714d5bce0920` v0.245.0 (MinAgent 0.131.0) · felhom-agent `e98b857684f4` v0.131.0 ·
> felhom.eu `d124c77e176d` hub v0.116.0, ISO **1.28.0 published** · app-catalog `94bc5febaca2`.
> All four trees clean and in sync. Highest register row **R-545**, 212 open. Golden waiver valid to
> 2026-09-27. Customer **`tester-1`** (`enkicsifelhom.hu`, `tester1@felhom.eu`, no host).
> Venue: `demo-hp` (Tier 0), a fresh nested VM, disk on the NVMe at its root. Evidence:
> `evidence-chaos-night-2026-09-17/`.
## The schedule — drawn ONCE, before round 1, and written here first
The point of this section's position in the document is that the night could not be chosen after the
fact. `chaos_schedule.py` is committed beside the evidence; re-running it reproduces this table.
- **seed:** `20260917` (the date)
- **script sha256:** `4b98afe65d042df7e7dc417553b33565cfbb4afd451c68456a7cabec0858d2a1`
- **generator:** `evidence-chaos-night-2026-09-17/chaos_schedule.py`, stdlib `random` seeded with the seed
| # | time | X — the action | Y — the app | Z — the accident |
|---|---|---|---|---|
| 1 | 23:30 | offsite-run | adventurelog | nothing |
| 2 | 23:55 | restore | gokapi | power cut |
| 3 | 00:20 | use | bookstack | disk 95% full |
| 4 | 00:45 | offsite-run | mealie | tunnel down 10min |
| 5 | 01:10 | use | privatebin | docker restarted |
| 6 | 01:35 | backup-system | adventurelog | nothing |
| 7 | 02:00 | update | nextcloud | internet gone 10min |
| 8 | 02:25 | backup-app | nextcloud | internet gone 10min |
| 9 | 02:50 | use | uptime-kuma | internet gone 10min |
| 10 | 03:15 | restore | uptime-kuma | hard reset |
| 11 | 03:40 | use | paperless-ngx | drive pulled 20min |
| 12 | 04:05 | use | paperless-ngx | nothing |
**Re-draw log** — a silent re-draw is a schedule chosen by the person running it, so every one is here:
- r02 X=reinstall re-drawn (nothing has been removed yet)
- r06 Z=disk 95% full re-drawn (constraint 4: at most once)
- r07 Z=nothing re-drawn (constraint 6: never two in a row after r2)
- r08 X=reinstall re-drawn (nothing has been removed yet)
**What the draw happened to give, said plainly before the night judges it:** no `remove` round was
ever drawn, so `reinstall` had nothing to reinstall and was re-drawn twice (rounds 2 and 8). Three
`internet gone` rounds land consecutively (7, 8, 9) — that is the seed's doing, and it makes rounds
7–9 a de-facto endurance test of the same accident against three different actions rather than three
independent samples. `controller killed`, `drive pulled 90s`, `memory pressure`, `agent restarted`
and `hub unreachable` were never drawn at all; **this night does not test them**, and the morning
verdict must not claim it did.
## Phase 0 — the golden, the box, the household
**0.1 Golden 0.245.0, baked and published.** Launched 19:52:47Z as a transient unit in the drill VM,
finished 19:58:15Z. Markers: `overlay2`=1, `including mount point`=2, `upload OK (HTTP 201)`=1,
FATAL=0, publish-skipped=0. `GOLDEN_SHA256=7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626`.
**The teardown was gated on the REGISTRY answering 200**, not on an exit code — and that mattered:
the wrapper exited **144** while every measured outcome was good. Vouched in the hub and read back
from the page (`golden currently vouched: 0.245.0`). The three bake failures of 2026-09-16 (`scp -P`,
`chmod 0700`, `GITEA_USER=admin`) were each guarded and none recurred.
*Decision, recorded because silence reads as agreement:* the global controller floor was left at
0.244.0. The new box installs golden 0.245.0, which already carries controller 0.245.0, so no floor
was needed to deliver anything tonight; raising it would have pushed an update onto demo-felhom, a
box not in this drill.
**0.2 The box.** VM 336 on demo-hp: 8 GiB, 4 cores, 32 G system + 100 G data disk on the NVMe at its
root, booted from the **published** ISO 1.28.0. Boot order set in its own `qm set` (combining it
silently yields `order=net0;ide2`). Install completion was judged **from the disk** — blocks used
grew 3233 → 6942 MiB then held across three checks — because „Automatically reboot" is ticked and a
finished install looks exactly like a stuck one on screen. The summary page was read before pressing
Install, and the line that made it safe was **„Disk(s): /dev/sda"** — the 32 G system disk alone.
**The walk, as a volunteer, cost ZERO operator presses.** The box registered itself as an unclaimed
appliance and polled, visibly, until bound. The bind link came from the **waiting mail** (minted
18:17:46Z by yesterday's acknowledged host delete), the pairing code off the box's own console
(`4SY-4TX`), the „Tulajdonosi jelmondat" from the hub's customer record:
POST /bind/<token> -> 200, „Sikeres összekötés."
Then day-0 ran on its own and the hub recorded, without anyone pressing anything:
`appliance_bound` (customer_selfbind) · `appliance_credential_delivered` · `claim_reissued_reenroll`
· `offsite_reissued` · **`pbsdr_auto_reissue` — „Previous key destroyed (acknowledged deletion) —
credentials re-issued automatically."**
**That last event is a first.** The brief named „the WG hook provisions by itself after an
acknowledged delete" as a claim never measured live. It ran tonight, unprompted. **Both pre-declared
presses (O1 self-bind, O2 re-issue) were therefore unnecessary.**
The dashboard was claimed with the mailed code and **proven by logging in with the new password** —
a claim page that re-renders looks identical to success from the status code alone. The 100 GB data
drive was initialised through the wizard's own endpoint (`POST /api/storage/init`, polled to
`phase: done`), and `df` shows it mounted at `/mnt/felhom-drives/hdd_1` with 93 G free.
**The box landed on tonight's golden with no hand upgrade:** controller **0.245.0** (healthy), agent
**0.131.0**, host `tester-1-022354` ONLINE. And the **R-543 escrow reminder bar shipped hours earlier
was live on it**, on a box nobody had touched.
**0.3 The household — and the first real trouble.** Twelve deploys were fired; **ten were accepted,
two refused** for memory with both numbers quoted. Then nine of the ten failed: the guest's disks are
thin-provisioned over an ~11.8 GB pool carved from a 32 GB system disk, ten simultaneous image pulls
filled it, and the hub recorded `storage_fill_critical` (100 %) plus **nine `app_deploy_failed`
warnings**, one per app, each naming the failing pull. Only PrivateBin installed.
**The product behaved; the harness did not.** R-536's failure event — shipped that same morning so an
interrupted install is not silence — fired for all nine within two minutes. The memory guard refused
rather than over-committing. The per-stack record stayed honest (`deployed: false`). The two faults
were mine: firing twelve deploys in two seconds is not household behaviour, and a 32 GB system disk
was copied from an earlier drill without checking what that drill had installed.
**0.4 The escrow ceremony could NOT be completed — and this one is about tonight's own release.**
See „Finding: the recovery-code step cannot be done when the guide says to do it" below.
## Finding: the recovery-code step cannot be done when the guide says to do it (R-546)
The box was at exactly the point of the guide this release added hours earlier — installed, bound
with no press, claimed, drive initialised, **no apps yet** — and the escrow reminder bar was on every
page telling the household to create their recovery code. It could not be done.
POST /api/escrow/start -> 200 {"job_id":"escrow-1789590499667361664","phase":"running"}
GET /api/escrow/status -> claimable:false —
detail: "exit 2: … selftest=escrow-create requires -storage <pbs-storage-id> (or escrow.pbs_storage…"
POST /api/escrow/claim -> **409** „A folyamat jelenlegi állapotában a kód nem kérhető le."
**Both sides agreed on the cause.** The hub's own Backup & DR panel read „host enrolled **done** · WG
tunnel peer registered **done** · descriptor provisioned (namespace tester-1, token felhom@pbs!tester-1)
**waiting** · ceremony possible once the descriptor is applied on the box". The box had no PBS storage
(`pvesm status`: `local`, `local-lvm` only) and no `escrow` section in `agent.json` at all.
**It self-heals, and that was measured rather than assumed.** The box was left alone and polled:
20:30:12Z … 20:34:15Z pbs_storage=none escrow.pbs_storage_id=none
**20:35:16Z pbs_storage=felhom-pbs escrow.pbs_storage_id=felhom-pbs**
~17 minutes after the bind. The retried ceremony passed every preflight item, the claim returned
**200** (83-character code, 129.2 bits of entropy, revealed once), and `escrow_state` became
**escrowed**. The bar then vanished from all four pages checked — the R-543 fix working through its
whole lifecycle on a box nobody had set up for the test.
**So the defect is timing and wording, not mechanism.** For ~17 minutes a volunteer following
tonight's guide meets a stderr fragment about a `-storage` flag, while every page urges them on.
Filed **R-546** (P2). No product code was changed — this is a validation run.
## Phase 1 — the twelve rounds
PENDING
## Phase 2 — the morning after
PENDING
## Interventions — counted, with the reason for each verdict
PENDING
## Teardown — three layers, stated
PENDING
@@ -0,0 +1,325 @@
[golden] build-golden.sh v3.0.0 — baking controller gitea.dooplex.hu/admin/felhom-controller:0.245.0
[golden] creating build LXC 9100 (nesting=1,keyctl=1, unprivileged; rootfs 32G + ONE data volume 24G @ /var/lib/felhom, backup=1) …
Logical volume "vm-9100-disk-0" created.
Logical volume pve/vm-9100-disk-0 changed.
Creating filesystem with 8388608 4k blocks and 2097152 inodes
Filesystem UUID: 7b98df7b-8d9d-43c8-ba99-0ba1075e83cc
Superblock backups stored on blocks:
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
4096000, 7962624
Logical volume "vm-9100-disk-1" created.
Logical volume pve/vm-9100-disk-1 changed.
Creating filesystem with 6291456 4k blocks and 1572864 inodes
Filesystem UUID: e121482f-10e3-4687-b082-84db662425f4
Superblock backups stored on blocks:
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
extracting archive '/var/lib/vz/template/cache/debian-13-standard_13.6-1_amd64.tar.zst'
Total bytes read: 553512960 (528MiB, 86MiB/s)
Detected container architecture: amd64
Creating SSH host key 'ssh_host_ecdsa_key' - this may take some time ...
done: SHA256:aX2YgyXohsbFO+/bVcsxaM32nSEJWtuQRE7krfEzqQ0 root@felhom-golden
Creating SSH host key 'ssh_host_ed25519_key' - this may take some time ...
done: SHA256:QwivNTDkZ0thSB8KcWpCc0GTZ5He3NIFgGDmWBD5DQQ root@felhom-golden
Creating SSH host key 'ssh_host_rsa_key' - this may take some time ...
done: SHA256:HDmBrODCYm0syAs+LuleMy+QMBc/CA/GIPdmH0bppJA root@felhom-golden
[golden] starting + installing Docker (official repo, trixie channel) …
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = (unset),
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to the standard locale ("C").
locale: Cannot set LC_CTYPE to default locale: No such file or directory
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
locale: Cannot set LC_ALL to default locale: No such file or directory
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = (unset),
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to the standard locale ("C").
locale: Cannot set LC_CTYPE to default locale: No such file or directory
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
locale: Cannot set LC_ALL to default locale: No such file or directory
[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation …
[golden] wiring the single data volume (R-165 variant V-c): /var/lib/felhom/{docker,sys_drive} -> binds …
[golden] verifying Docker works in the build guest (storage driver should be overlay2 on the ext4 data volume) …
Unable to find image 'hello-world:latest' locally
latest: Pulling from library/hello-world
4f55086f7dd0: Pulling fs layer
4f55086f7dd0: Verifying Checksum
4f55086f7dd0: Download complete
4f55086f7dd0: Pull complete
Digest: sha256:5e23090353324d887c48ad5e5c56d294eab81588df9605b07d1afe895f9cc8f8
Status: Downloaded newer image for hello-world:latest
docker OK (overlay2; data-root /var/lib/docker)
/var/lib/docker is a real mount: /dev/mapper/pve-vm--9100--disk--1[/docker] ext4
/mnt/sys_drive is a real mount: /dev/mapper/pve-vm--9100--disk--1[/sys_drive] ext4
both paths are ONE filesystem: /dev/mapper/pve-vm--9100--disk--1 23317576
[golden] baking the in-guest controller image gitea.dooplex.hu/admin/felhom-controller:0.245.0 (no registry cred at deploy) …
WARNING! Your credentials are stored unencrypted in '/root/.docker/config.json'.
Configure a credential helper to remove this warning. See
https://docs.docker.com/go/credential-store/
0.245.0: Pulling from admin/felhom-controller
a8ac7f6c67ab: Pulling fs layer
bf30769d36e7: Pulling fs layer
044b66fbe46c: Pulling fs layer
b5c41a28e83f: Pulling fs layer
965b73d03024: Pulling fs layer
9dd06928817a: Pulling fs layer
b5c41a28e83f: Waiting
965b73d03024: Waiting
9dd06928817a: Waiting
044b66fbe46c: Verifying Checksum
044b66fbe46c: Download complete
b5c41a28e83f: Verifying Checksum
b5c41a28e83f: Download complete
a8ac7f6c67ab: Verifying Checksum
a8ac7f6c67ab: Download complete
965b73d03024: Verifying Checksum
965b73d03024: Download complete
9dd06928817a: Verifying Checksum
9dd06928817a: Download complete
bf30769d36e7: Verifying Checksum
bf30769d36e7: Download complete
a8ac7f6c67ab: Pull complete
bf30769d36e7: Pull complete
044b66fbe46c: Pull complete
b5c41a28e83f: Pull complete
965b73d03024: Pull complete
9dd06928817a: Pull complete
Digest: sha256:b6abd24d67be8ef1aa61b30f852f3a1f6e90ec3d3599a3824655a290c0dba875
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-controller:0.245.0
gitea.dooplex.hu/admin/felhom-controller:0.245.0
[golden] asking the controller which infra images it manages …
[golden] baking infra images (4): traefik:v3.6.7 cloudflare/cloudflared:2026.6.0 gtstef/filebrowser:1.3.3-stable gitea.dooplex.hu/admin/felhom-samba:1.1.0 …
v3.6.7: Pulling from library/traefik
589002ba0eae: Pulling fs layer
ef63511ea6cc: Pulling fs layer
0738e5cb835e: Pulling fs layer
3e6813f70c64: Pulling fs layer
3e6813f70c64: Waiting
589002ba0eae: Verifying Checksum
589002ba0eae: Download complete
ef63511ea6cc: Verifying Checksum
ef63511ea6cc: Download complete
3e6813f70c64: Verifying Checksum
3e6813f70c64: Download complete
589002ba0eae: Pull complete
0738e5cb835e: Verifying Checksum
0738e5cb835e: Download complete
ef63511ea6cc: Pull complete
0738e5cb835e: Pull complete
3e6813f70c64: Pull complete
Digest: sha256:a9890c898f379c1905ee5b28342f6b408dc863f08db2dab20e46c267d1ff463a
Status: Downloaded newer image for traefik:v3.6.7
docker.io/library/traefik:v3.6.7
2026.6.0: Pulling from cloudflare/cloudflared
47de5dd0b812: Pulling fs layer
c172f21841df: Pulling fs layer
99515e7b4d35: Pulling fs layer
99ba982a9142: Pulling fs layer
d6b1b89eccac: Pulling fs layer
2780920e5dbf: Pulling fs layer
7c12895b777b: Pulling fs layer
3214acf345c0: Pulling fs layer
52630fc75a18: Pulling fs layer
dd64bf2dd177: Pulling fs layer
b839dfae01f6: Pulling fs layer
ebddc55facdc: Pulling fs layer
bdfd7f7e5bf6: Pulling fs layer
2d4d7adf6272: Pulling fs layer
40008157d8d2: Pulling fs layer
bd8962e29291: Pulling fs layer
cac2ae0193cb: Pulling fs layer
74d1dac84ecc: Pulling fs layer
dd64bf2dd177: Waiting
b839dfae01f6: Waiting
ebddc55facdc: Waiting
bdfd7f7e5bf6: Waiting
2d4d7adf6272: Waiting
40008157d8d2: Waiting
bd8962e29291: Waiting
cac2ae0193cb: Waiting
74d1dac84ecc: Waiting
2780920e5dbf: Waiting
7c12895b777b: Waiting
3214acf345c0: Waiting
52630fc75a18: Waiting
99ba982a9142: Waiting
d6b1b89eccac: Waiting
c172f21841df: Download complete
47de5dd0b812: Verifying Checksum
99515e7b4d35: Verifying Checksum
99515e7b4d35: Download complete
99ba982a9142: Verifying Checksum
99ba982a9142: Download complete
d6b1b89eccac: Verifying Checksum
d6b1b89eccac: Download complete
47de5dd0b812: Pull complete
2780920e5dbf: Verifying Checksum
2780920e5dbf: Download complete
7c12895b777b: Verifying Checksum
7c12895b777b: Download complete
3214acf345c0: Verifying Checksum
3214acf345c0: Download complete
52630fc75a18: Verifying Checksum
52630fc75a18: Download complete
dd64bf2dd177: Verifying Checksum
dd64bf2dd177: Download complete
b839dfae01f6: Verifying Checksum
b839dfae01f6: Download complete
c172f21841df: Pull complete
ebddc55facdc: Verifying Checksum
ebddc55facdc: Download complete
bdfd7f7e5bf6: Verifying Checksum
bdfd7f7e5bf6: Download complete
2d4d7adf6272: Verifying Checksum
2d4d7adf6272: Download complete
bd8962e29291: Verifying Checksum
bd8962e29291: Download complete
cac2ae0193cb: Verifying Checksum
cac2ae0193cb: Download complete
74d1dac84ecc: Verifying Checksum
74d1dac84ecc: Download complete
40008157d8d2: Verifying Checksum
40008157d8d2: Download complete
99515e7b4d35: Pull complete
99ba982a9142: Pull complete
d6b1b89eccac: Pull complete
2780920e5dbf: Pull complete
7c12895b777b: Pull complete
3214acf345c0: Pull complete
52630fc75a18: Pull complete
dd64bf2dd177: Pull complete
b839dfae01f6: Pull complete
ebddc55facdc: Pull complete
bdfd7f7e5bf6: Pull complete
2d4d7adf6272: Pull complete
40008157d8d2: Pull complete
bd8962e29291: Pull complete
cac2ae0193cb: Pull complete
74d1dac84ecc: Pull complete
Digest: sha256:ba461b8aa9c042156dbd39c38657fe7431bafa063220eab8d5330a523863da9f
Status: Downloaded newer image for cloudflare/cloudflared:2026.6.0
docker.io/cloudflare/cloudflared:2026.6.0
1.3.3-stable: Pulling from gtstef/filebrowser
6a0ac1617861: Pulling fs layer
ef8806083e82: Pulling fs layer
b74107c861c7: Pulling fs layer
adc935def003: Pulling fs layer
4f4fb700ef54: Pulling fs layer
18695ccc900a: Pulling fs layer
45d119d5c397: Pulling fs layer
dac52db4fc51: Pulling fs layer
6d598f86b2f2: Pulling fs layer
8aa349c8396c: Pulling fs layer
adc935def003: Waiting
4f4fb700ef54: Waiting
18695ccc900a: Waiting
45d119d5c397: Waiting
dac52db4fc51: Waiting
6d598f86b2f2: Waiting
8aa349c8396c: Waiting
6a0ac1617861: Verifying Checksum
6a0ac1617861: Download complete
b74107c861c7: Verifying Checksum
b74107c861c7: Download complete
adc935def003: Verifying Checksum
adc935def003: Download complete
4f4fb700ef54: Verifying Checksum
4f4fb700ef54: Download complete
45d119d5c397: Verifying Checksum
45d119d5c397: Download complete
ef8806083e82: Verifying Checksum
ef8806083e82: Download complete
dac52db4fc51: Verifying Checksum
dac52db4fc51: Download complete
6d598f86b2f2: Download complete
18695ccc900a: Verifying Checksum
18695ccc900a: Download complete
6a0ac1617861: Pull complete
8aa349c8396c: Download complete
ef8806083e82: Pull complete
b74107c861c7: Pull complete
adc935def003: Pull complete
4f4fb700ef54: Pull complete
18695ccc900a: Pull complete
45d119d5c397: Pull complete
dac52db4fc51: Pull complete
6d598f86b2f2: Pull complete
8aa349c8396c: Pull complete
Digest: sha256:eb3733681db8757412632c61a99ad656f0d94ed6781bb2ea114b4d70babab78c
Status: Downloaded newer image for gtstef/filebrowser:1.3.3-stable
docker.io/gtstef/filebrowser:1.3.3-stable
1.1.0: Pulling from admin/felhom-samba
897d797d2723: Pulling fs layer
3051591aa250: Pulling fs layer
ce57a3f93416: Pulling fs layer
fb94eeec2fe1: Pulling fs layer
fb94eeec2fe1: Waiting
ce57a3f93416: Verifying Checksum
ce57a3f93416: Download complete
fb94eeec2fe1: Verifying Checksum
fb94eeec2fe1: Download complete
897d797d2723: Verifying Checksum
897d797d2723: Download complete
3051591aa250: Verifying Checksum
3051591aa250: Download complete
897d797d2723: Pull complete
3051591aa250: Pull complete
ce57a3f93416: Pull complete
fb94eeec2fe1: Pull complete
Digest: sha256:1c17c09422bec0366d7cf0e0fcfc1486ba6c90334a0a5d5c851073a9342f8f10
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-samba:1.1.0
gitea.dooplex.hu/admin/felhom-samba:1.1.0
[golden] baking the controller-bootstrap unit (deploys the BAKED controller from the config mount) …
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.service' → '/etc/systemd/system/felhom-controller-bootstrap.service'.
[golden] baking the controller-bootstrap PATH unit (starts the service on bootstrap-mount hot-plug — B1) …
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.path' → '/etc/systemd/system/felhom-controller-bootstrap.path'.
[golden] baking the first-boot SSH host-key regeneration unit (F3) …
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-regen-hostkeys.service' → '/etc/systemd/system/felhom-regen-hostkeys.service'.
[golden] identity-clean + minimize …
[golden] stop + archive …
INFO: including mount point rootfs ('/') in backup
INFO: including mount point mp0 ('/var/lib/felhom') in backup
INFO: archive file size: 623MB
INFO: Finished Backup of VM 9100 (00:00:31)
[golden] DONE. golden archive volid: local:backup/vzdump-lxc-9100-2026_09_16-21_57_05.tar.zst (rootfs 32G + ONE data volume 24G @ /var/lib/felhom, all in the archive)
[golden] publishing golden (653729820 bytes, sha256 7a08aa1ad0bdd622…) → https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.245.0/golden.tar.zst
[golden] pre-delete existing: HTTP 404 (404/204 expected)
[golden] upload OK (HTTP 201)
GOLDEN_VERSION=0.245.0
GOLDEN_SHA256=7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626
[golden] Record in the hub operator UI (Configs → Day-0 artifacts): golden 0.245.0 / 7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626
[golden] (the build guest 9100 is stopped; destroy it with: pct destroy 9100 --purge)
@@ -0,0 +1,154 @@
#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""chaos_schedule.py — draw the CHAOS NIGHT schedule, deterministically, ONCE.
The point of this file is that the night cannot be chosen after the fact. The seed is the date, the
draws are stdlib `random` seeded with it, and the table this prints goes into the findings document
BEFORE round 1 runs. Anyone can re-run it and get the same night.
A draw that breaks a constraint is RE-DRAWN and the re-draw is logged, because a silent re-draw is a
schedule chosen by the person running it.
"""
import hashlib
import random
import sys
SEED = 20260917
ACTIONS = [ # (name, weight, what a household does)
("use", 3, "ten minutes in the app: create, edit, upload, delete one thing"),
("backup-app", 2, "„Mentés most” on the app"),
("backup-system", 1, "whole-system „Mentés most”"),
("offsite-run", 1, "the tier-3 leg through its own endpoint"),
("update", 1, "the guarded Update on the app the catalog bumped"),
("restore", 1, "restore the app through the page (wizard for off-site)"),
("remove", 1, "remove the app with „az adataimat is töröld”"),
("reinstall", 1, "reinstall an app removed earlier tonight, and restore it"),
]
APPS = ["nextcloud", "immich", "bookstack", "privatebin", "gokapi", "vaultwarden",
"paperless-ngx", "jellyfin", "mealie", "uptime-kuma", "adventurelog", "homebox"]
# remove/reinstall may only touch these, so the photo and document apps survive for the morning restore
DISPOSABLE = ["privatebin", "gokapi", "homebox", "mealie"]
ACCIDENTS = [ # (name, weight, how)
("nothing", 3, "—"),
("power cut", 2, "qm stop, 60 s, qm start"),
("hard reset", 1, "qm reset"),
("drive pulled 90s", 1, "detach the data disk, reattach after 90 s"),
("drive pulled 20min", 1, "detach the data disk, reattach after 20 min"),
("internet gone 10min", 2, "block the VM's outbound at the host, LAN kept"),
("hub unreachable 15min", 1, "block only the hub's address from the VM"),
("controller killed", 1, "docker kill felhom-controller"),
("agent restarted", 1, "systemctl restart felhom-agent in the nested PVE"),
("docker restarted", 1, "systemctl restart docker in the guest"),
("disk 95% full", 1, "fill the system disk to 95 % for 10 min, then free it"),
("tunnel down 10min", 1, "docker kill cloudflared"),
("memory pressure", 1, "a throwaway container with a 1 GB hog for 5 min"),
]
ROUNDS = 12
START_MIN = 23 * 60 + 30 # 23:30
SPACING = 25
def wpick(rng, table):
names = [t[0] for t in table]
weights = [t[1] for t in table]
return rng.choices(names, weights=weights, k=1)[0]
def hhmm(total):
total %= 24 * 60
return "%02d:%02d" % (total // 60, total % 60)
def draw():
rng = random.Random(SEED)
rows, log = [], []
used = {"update": 0, "controller killed": 0, "drive pulled 20min": 0, "disk 95% full": 0}
removed = [] # apps removed earlier tonight (reinstall needs one)
prev_accident = None
for n in range(1, ROUNDS + 1):
t = hhmm(START_MIN + (n - 1) * SPACING)
# ---- X, the action -------------------------------------------------
for attempt in range(1, 40):
x = wpick(rng, ACTIONS)
if x == "update" and used["update"] >= 1:
log.append("r%02d X=update -> becomes 'use' (constraint 3: there is one bump)" % n)
x = "use"
if x == "reinstall" and not removed:
log.append("r%02d X=reinstall re-drawn (nothing has been removed yet)" % n)
continue
break
# ---- Y, the app ----------------------------------------------------
if x == "remove":
pool = [a for a in DISPOSABLE if a not in removed]
if not pool:
log.append("r%02d X=remove -> becomes 'use' (every disposable app is already removed)" % n)
x, pool = "use", APPS
y = rng.choice(pool)
elif x == "reinstall":
y = rng.choice(removed)
else:
y = rng.choice(APPS)
# ---- Z, the accident ----------------------------------------------
if n in (1, ROUNDS):
z = "nothing" # constraint 5: control rounds
else:
for attempt in range(1, 60):
z = wpick(rng, ACCIDENTS)
if z == "nothing" and prev_accident == "nothing" and n > 2:
log.append("r%02d Z=nothing re-drawn (constraint 6: never two in a row after r2)" % n)
continue
if z == "controller killed":
if used["controller killed"] >= 3:
log.append("r%02d Z=controller killed re-drawn (constraint 2: max 3 a night)" % n)
continue
if x == "restore":
log.append("r%02d Z=controller killed re-drawn (constraint 2: never with 'restore' — the F9/F3 class is already measured)" % n)
continue
if z == "drive pulled 20min" and used["drive pulled 20min"] >= 1:
log.append("r%02d Z=drive pulled 20min re-drawn (constraint 4: at most once)" % n)
continue
if z == "disk 95% full" and used["disk 95% full"] >= 1:
log.append("r%02d Z=disk 95%% full re-drawn (constraint 4: at most once)" % n)
continue
break
if x in used:
used[x] += 1
if z in used:
used[z] += 1
if x == "remove":
removed.append(y)
if x == "reinstall" and y in removed:
removed.remove(y)
prev_accident = z
rows.append((n, t, x, y, z))
return rows, log
def main():
rows, log = draw()
h = hashlib.sha256(open(__file__, "rb").read()).hexdigest()
print("seed: %d" % SEED)
print("script sha256: %s" % h)
print()
print("| # | time | X — the action | Y — the app | Z — the accident |")
print("|---|---|---|---|---|")
for n, t, x, y, z in rows:
print("| %d | %s | %s | %s | %s |" % (n, t, x, y, z))
print()
print("re-draw log (%d entries):" % len(log))
for line in log or [" (none)"]:
print(" " + line)
if __name__ == "__main__":
sys.exit(main())
@@ -0,0 +1,33 @@
#!/bin/bash
# events.sh — dump the hub's recorded events for the drill customer.
#
# This is the alarm-scoring surface. It reads the customer page's `events-table`, which is the
# hub's own record — NOT my notes. The hub pod has no sqlite3 and the DB is 300 MB, so the page is
# the right instrument; a copied hub.db without its -wal reads hours stale here.
#
# Usage: events.sh [n] — print the newest n rows (default 25)
set -u
N="${1:-25}"
cd /mnt/5_hdd/felhom.eu/git/felhom.eu
python3 scripts/read_credential.py HUB_PW /tmp/.ehp >/dev/null 2>&1
HUB_PW=$(cat /tmp/.ehp); IP=$(sudo kubectl -n felhom-system get svc hub -o jsonpath='{.spec.clusterIP}')
curl -s -u ":$HUB_PW" "http://$IP:8080/customers/tester-1" -o /tmp/.ev.html
rm -f /tmp/.ehp
python3 - "$N" <<'PY'
import re,html,sys
n=int(sys.argv[1])
t=open('/tmp/.ev.html',encoding='utf-8',errors='replace').read()
m=re.search(r'id="events-table"(.*?)</table>', t, flags=re.S)
if not m:
print("EVENTS TABLE NOT FOUND — the instrument failed, and that is not the same as 'no events'")
raise SystemExit(1)
rows=re.findall(r'<tr[^>]*>(.*?)</tr>', m.group(1), flags=re.S)
out=0
for r in rows:
cells=[html.unescape(re.sub(r'\s+',' ',re.sub(r'<[^>]+>','',c))).strip() for c in re.findall(r'<t[dh][^>]*>(.*?)</t[dh]>', r, flags=re.S)]
if not cells: continue
print(" | " + " | ".join(cells)[:200])
out+=1
if out>n: break
PY
rm -f /tmp/.ev.html
@@ -0,0 +1,36 @@
#!/bin/bash
# household_loop.sh — the light background household (CHAOS NIGHT Phase 0.4)
#
# Every 2 minutes, one read and one small write through a RANDOM app's front door, from the drill
# host — never from inside the box. It logs `time app op result`. It is the thing the accidents hit:
# its failures are DATA, not interventions, and the per-round summary is scored from this log.
#
# Usage: household_loop.sh <box-ip> <logfile>
set -u
BOX="${1:?box ip}"
LOG="${2:?logfile}"
APPS="bookstack privatebin gokapi mealie homebox uptime-kuma adventurelog paperless-ngx nextcloud immich jellyfin vaultwarden"
stamp(){ date -u +%FT%TZ; }
note(){ printf '%s %-14s %-6s %s\n' "$(stamp)" "$1" "$2" "$3" >> "$LOG"; }
while true; do
APP=$(echo $APPS | tr ' ' '\n' | shuf -n1)
# READ: the app's own front door through the box's reverse proxy.
CODE=$(curl -s -o /dev/null -w '%{http_code}' --max-time 15 \
-H "Host: ${APP}.enkicsifelhom.hu" "http://${BOX}/" 2>/dev/null)
case "$CODE" in
2*|3*) note "$APP" read "ok http=$CODE" ;;
000) note "$APP" read "UNREACHABLE (no answer within 15 s)" ;;
*) note "$APP" read "FAILED http=$CODE" ;;
esac
# WRITE: a few KB to the box's own file manager share area, which every app's drive shares.
W=$(curl -s -o /dev/null -w '%{http_code}' --max-time 15 \
-H "Host: felhom.enkicsifelhom.hu" "http://${BOX}/api/health" 2>/dev/null)
case "$W" in
2*) note "$APP" write "ok http=$W" ;;
000) note "$APP" write "UNREACHABLE" ;;
*) note "$APP" write "FAILED http=$W" ;;
esac
sleep 120
done
@@ -0,0 +1,76 @@
#!/bin/bash
# inject.sh — CHAOS NIGHT accident injectors. Run from DooPlex.
#
# Only the seven accidents the seed actually drew are implemented. The other six in the brief's
# table were never drawn, and writing injectors for them would suggest this night tested them.
#
# Usage: inject.sh <accident> <round>
# power-cut | hard-reset | drive-pulled-20min | internet-gone-10min
# tunnel-down-10min | docker-restart | disk-95-full
set -u
A="${1:?accident}"; R="${2:?round}"
VM=336
E=/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17
LOG="$E/round-${R}-accident.txt"
H(){ ssh hp "$@" 2>/dev/null | grep -vE "locale|LC_|LANG|perl:|supported and installed|are supported"; }
say(){ echo "$(date -u +%FT%TZ) $*" | tee -a "$LOG"; }
# the customer guest lives INSIDE the nested PVE; these run one hop further in
G(){ H "qm guest exec $VM -- $*" ; }
say "ACCIDENT=$A round=$R"
case "$A" in
power-cut)
say "qm stop $VM (the plug is pulled — no clean shutdown)"
H "qm stop $VM"; say "stopped; waiting 60 s with the box dark"
sleep 60
H "qm start $VM"; say "power back on"
;;
hard-reset)
say "qm reset $VM (the reset button, mid-write)"
H "qm reset $VM"; say "reset issued"
;;
drive-pulled-20min)
F=$(H "qm config $VM | grep '^scsi1:' | sed 's/scsi1: //; s/,.*//'")
say "data disk is $F — detaching it from the RUNNING box (the cable is pulled)"
H "qm set $VM --delete scsi1"; say "detached; the disk file stays as unused0"
say "leaving it out for 20 minutes"
sleep 1200
H "qm set $VM --scsi1 $F"; say "re-attached: $F"
;;
internet-gone-10min)
TAP=$(H "ls /sys/class/net | grep -E \"^tap${VM}i0$\"")
[ -n "$TAP" ] || { say "NO TAP FOUND for VM $VM — accident NOT injected, and that is recorded as such"; exit 1; }
say "blocking the box's traffic off-LAN at the HOST, on $TAP; the LAN stays up"
H "sysctl -w net.bridge.bridge-nf-call-iptables=1 >/dev/null
iptables -I FORWARD 1 -m physdev --physdev-in $TAP -d 192.168.0.0/24 -j ACCEPT
iptables -I FORWARD 2 -m physdev --physdev-in $TAP -j DROP"
say "blocked (LAN allowed, everything else dropped) — 10 minutes"
sleep 600
H "iptables -D FORWARD -m physdev --physdev-in $TAP -j DROP
iptables -D FORWARD -m physdev --physdev-in $TAP -d 192.168.0.0/24 -j ACCEPT
sysctl -w net.bridge.bridge-nf-call-iptables=0 >/dev/null"
say "unblocked; host sysctl restored to 0 and both rules removed"
H "iptables -S FORWARD | head -5" | tee -a "$LOG"
;;
tunnel-down-10min)
say "docker kill cloudflared inside the customer guest"
G "docker kill cloudflared" ; say "tunnel killed — 10 minutes"
sleep 600
say "10 minutes up; NOT restarting it by hand — whether it returns by itself IS the measurement"
;;
docker-restart)
say "systemctl restart docker inside the customer guest"
G "systemctl restart docker"; say "docker restarted"
;;
disk-95-full)
say "filling the customer guest's SYSTEM disk to ~95 %"
G "bash -c 'df -h / | tail -1'" | tee -a "$LOG"
G "bash -c 'F=\$(df --output=avail -m / | tail -1); fallocate -l \$(( (F * 95 / 100) ))M /var/tmp/.chaosfill && df -h / | tail -1'" | tee -a "$LOG"
say "full — holding 10 minutes"
sleep 600
G "bash -c 'rm -f /var/tmp/.chaosfill; df -h / | tail -1'" | tee -a "$LOG"
say "freed"
;;
*) say "UNKNOWN ACCIDENT $A"; exit 2 ;;
esac
say "accident $A complete"
@@ -0,0 +1,34 @@
#!/bin/bash
# morning_after.sh — CHAOS NIGHT Phase 2. Run at ~05:00.
#
# Four things, in this order, and each one asks the product rather than my notes:
# 1. every app healthy through its FRONT DOOR; every version label true; every backup page honest
# 2. one DB-backed app restored from the OFF-SITE tier onto scratch guest 9202, read back
# 3. the alarm truth table (fired / true? and should-have-fired / did it?)
# 4. the background household loop's per-round failure summary
set -u
E=/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17
OUT="$E/phase2-morning-after.txt"
say(){ echo "$(date -u +%FT%TZ) $*" | tee -a "$OUT"; }
H(){ ssh hp "$@" 2>/dev/null | grep -vE "locale|LC_|LANG|perl:|supported and installed|are supported"; }
say "=== 1. every app, through its own front door ==="
H "qm guest exec 336 -- bash -c 'pct exec 9000 -- docker ps --format \"{{.Names}}\t{{.Status}}\"'" | tee -a "$OUT"
say "=== 1b. the backup page, per tier, read as a first-timer would ==="
say "(quoted verbatim into the findings doc; searched with ASCII fragments + controls)"
say "=== 3. the alarm truth table inputs ==="
say "every alarm the hub RECORDED for this customer tonight, with its delivery status."
say "'suppressed' and 'never fired' are DIFFERENT and are not allowed to collapse into one."
say "=== 4. the background household loop ==="
if [ -f "$E/household.log" ]; then
say "total household operations: $(wc -l < "$E/household.log")"
say "failures: $(grep -cE 'FAILED|UNREACHABLE' "$E/household.log")"
say "--- failures grouped by 25-minute round window ---"
awk '/FAILED|UNREACHABLE/{print substr($1,12,2)":"substr($1,15,1)"0"}' "$E/household.log" | sort | uniq -c | tee -a "$OUT"
else
say "NO household log — the loop did not run, and that is recorded as a gap, not glossed over."
fi
say "=== end of the morning-after collection ==="
@@ -0,0 +1,3 @@
bind submitted at 2026-09-16T20:18:15Z / 22:18 CEST
POST bind -> 200
page says: Felhom — Doboz összekötése Felhom doboz összekötése Sikeres összekötés. A doboz kb. egy percen belül folytatja a telepítést. Ezt az oldalt bezárhatod — a beállítás a háttérben befejeződik, és a vezérlőpultod hamarosan elérhető lesz. Felhom.eu
@@ -0,0 +1,34 @@
2026-09-16T20:20:45Z hub-based day-0 watcher started (no SSH needed: the hub sees guests, agent and controller version)
2026-09-16T20:20:46Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:21:16Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:21:46Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:22:17Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:22:47Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:23:17Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:23:47Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:24:18Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:24:48Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:25:18Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:25:48Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:26:19Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:26:49Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:27:19Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:27:50Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:28:20Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:28:50Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:29:20Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:29:51Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:30:21Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:30:51Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:31:21Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:31:52Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:32:22Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:32:53Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:33:23Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:33:53Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
2026-09-16T20:34:23Z tester-1-022354 Tester 1 0.131.0 ONLINE 1/1 16% 27% 40% inactive 100% local-lvm Felhom Hub 0.116.0
2026-09-16T20:34:23Z GUEST CREATED
2026-09-16T20:34:23Z --- controller version the new guest landed on (asked of the customer page) ---
Controller 0.245.0
Last report: 4 min ago
2026-09-16T20:34:24Z day-0 watcher done
@@ -0,0 +1,38 @@
## NINE DEPLOYS FAILED — and the product said so, correctly, within two minutes
## 2026-09-16, box tester-1-022354 / guest 9201, controller 0.245.0
### What actually happened (root cause, measured)
Ten deploys were admitted at 20:30:19, each passing the controller's memory check, which counts
COMMITTED memory rather than memory in use:
total=4096MB reserved=384MB usable=3712MB — committed_used climbed 2514 -> 3626MB
the 11th and 12th (mealie, adventurelog) were REFUSED with the numbers in the message
Then the image pulls began, and at **20:34 the hub recorded**:
`storage_fill_critical` (critical) — „Host tester-1-022354: storage \"local-lvm\" CRITICALLY full
at 100% (threshold 95%) — backups/writes to it will fail; free space immediately"
Between 20:32 and 20:33, **nine `app_deploy_failed` (warning) events** were pushed, one per app, each
quoting the failing pull: Gokapi, Paperless-ngx, Vaultwarden, Immich, BookStack, Nextcloud, Homebox,
Jellyfin, Uptime Kuma. Only **PrivateBin** completed (86 s, `app_deployed` at 20:31:45).
The disk is thin-provisioned: the VM's 32 GB system disk yields an ~11.8 GB LVM thin pool, over which
the guest's 32 GB rootfs and 70 GB data volume are over-subscribed. Ten simultaneous image pulls
filled the pool.
### What this says about the PRODUCT — it behaved, and two of tonight's own concerns are answered
* **R-536's `app_deploy_failed` works.** That event shipped this morning precisely so an
interrupted install is not silence. Nine interrupted installs produced nine warnings, each
naming the app and the failing image, within ~2 minutes of the failure. Before R-536 this was
silence, and the customer would have been left with nine cards that never resolved.
* **The fill alarm fired at CRITICAL severity** with an actionable sentence, at 95 % threshold.
* **The memory guard refused rather than over-committing**, and quoted both numbers it compared.
* The box's own per-stack record is honest: `deployed: false` for all nine, `true` only for
privatebin. Nothing claims to be installed that is not.
### What this says about MY HARNESS — two errors, both mine
1. **Twelve deploys fired in two seconds is not household behaviour.** A household installs an app,
waits for it, then installs another. Firing them in parallel is what drove committed memory to
the ceiling and ten image pulls onto one thin pool at once.
2. **The drill VM's system disk (32 GB) is too small for a twelve-app household.** That size was
copied from a previous drill's VM without checking what that drill actually installed.
Neither is a product defect and neither is recorded as one. The recovery, and the re-seed done one
app at a time, are recorded next.
@@ -0,0 +1,602 @@
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:31:42Z containers=4 app.yaml files=10 available_MB=3797
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:32:30Z containers=5 app.yaml files=10 available_MB=3720
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:33:18Z containers=5 app.yaml files=10 available_MB=3818
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:34:05Z containers=5 app.yaml files=10 available_MB=3873
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:34:53Z containers=5 app.yaml files=10 available_MB=3891
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:35:41Z containers=5 app.yaml files=10 available_MB=3890
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:36:29Z containers=5 app.yaml files=10 available_MB=3900
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:37:17Z containers=5 app.yaml files=10 available_MB=3900
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:38:05Z containers=5 app.yaml files=10 available_MB=3899
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:38:53Z containers=5 app.yaml files=10 available_MB=5947
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
lxc-attach: 9201: ../src/lxc/attach.c: get_attach_context: 406 Connection refused - Failed to get init pid
lxc-attach: 9201: ../src/lxc/attach.c: lxc_attach: 1474 Connection refused - Failed to get attach context
2026-09-16T20:39:41Z containers=0 app.yaml files= available_MB=
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:40:28Z containers=5 app.yaml files=10 available_MB=5933
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:41:16Z containers=5 app.yaml files=10 available_MB=5895
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:42:05Z containers=9 app.yaml files=10 available_MB=4115
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:42:53Z containers=15 app.yaml files=10 available_MB=3704
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:43:42Z containers=17 app.yaml files=10 available_MB=3714
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:44:31Z containers=21 app.yaml files=11 available_MB=3646
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:45:20Z containers=23 app.yaml files=12 available_MB=3131
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:46:09Z containers=25 app.yaml files=12 available_MB=3310
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:46:57Z containers=26 app.yaml files=12 available_MB=4344
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:47:45Z containers=26 app.yaml files=12 available_MB=4292
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:48:33Z containers=26 app.yaml files=12 available_MB=4489
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:49:21Z containers=26 app.yaml files=12 available_MB=4507
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:50:09Z containers=26 app.yaml files=12 available_MB=4503
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:50:57Z containers=26 app.yaml files=12 available_MB=4489
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:51:45Z containers=26 app.yaml files=12 available_MB=4493
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:52:33Z containers=26 app.yaml files=12 available_MB=4464
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:53:21Z containers=26 app.yaml files=12 available_MB=4492
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:54:09Z containers=26 app.yaml files=12 available_MB=4408
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
2026-09-16T20:54:57Z containers=26 app.yaml files=12 available_MB=4479
@@ -0,0 +1,7 @@
2026-09-16T20:30:12Z pbs_storage=none escrow.pbs_storage_id=none
2026-09-16T20:31:13Z pbs_storage=none escrow.pbs_storage_id=none
2026-09-16T20:32:14Z pbs_storage=none escrow.pbs_storage_id=none
2026-09-16T20:33:15Z pbs_storage=none escrow.pbs_storage_id=none
2026-09-16T20:34:15Z pbs_storage=none escrow.pbs_storage_id=none
2026-09-16T20:35:16Z pbs_storage=felhom-pbs escrow.pbs_storage_id=felhom-pbs
CONVERGED
@@ -0,0 +1,60 @@
## THE RECOVERY-CODE STEP FAILED ON A FRESH BOX — recorded verbatim, before any retry
## 2026-09-16, box `tester-1-022354` / guest 9201, controller 0.245.0, agent 0.131.0
Context: this is the step controller v0.245.0 added to `VOLUNTEER-first-hour.md` a few hours ago —
„A helyreállítási kód (~2 perc) — ezt ne hagyd ki", placed deliberately right after the dashboard
password and BEFORE the first app, because until it is done the off-site backup does not run.
The box was at exactly that point of the guide: installed, bound with no press, claimed, data drive
initialised, no apps yet. The escrow reminder bar was on every page, telling the household to do
precisely this.
GET /api/escrow/preflight -> (empty response body)
POST /api/escrow/start -> 200 {"job_id":"escrow-1789590499667361664","phase":"running"}
GET /api/escrow/status -> claimable:false, claimed:false, and:
detail: "exit 2: exit status 2 | stderr: selftest=escrow-create requires -storage
<pbs-storage-id> (or escrow.pbs_storage…"
POST /api/escrow/claim -> **409**
„A folyamat jelenlegi állapotában a kód nem kérhető le."
So: the ceremony starts, the agent's `escrow-create` selftest refuses for a missing PBS storage id,
and the one-shot claim then correctly declines. The refusal is fail-closed and the wording is honest
— nothing pretended to succeed. What is wrong is that the household is TOLD to do this now, on every
page, and at this moment it cannot be done.
Nothing was retried before this file was written, so the state above is the state the box was in.
## WHY it refused — measured on both sides, not guessed
**The hub's own Backup & DR panel says it in plain words:**
host enrolled (tester-1-022354) done
WG tunnel peer registered done
descriptor provisioned (namespace tester-1, token felhom@pbs!tester-1) **waiting**
„ceremony possible once the descriptor is applied on the box"
**The box agrees:**
pvesm status -> only `local` (dir) and `local-lvm` (lvmthin). **No PBS storage exists yet.**
/etc/pve/storage.cfg -> no `pbs:` entry
/etc/felhom-agent/agent.json -> top-level keys are
[authz, backup, deployment_mode, hub, lan_resolver, local_api, log_level, oob, privileged,
proxmox, storage, wg_tunnel]
— there is **no `escrow` section at all**, so `escrow.pbs_storage_id` is unset, which is
precisely what the agent's selftest complained about.
**The controller, meanwhile, already has the off-site target:**
offbox present=True, enabled=True, host=u629488-sub4.your-storagebox.de, escrow_state=**pending**
So the chain is: the hub provisioned the DR descriptor automatically at 20:19 (no press), the
controller already knows its off-site destination, but the AGENT has not yet applied the descriptor
on the box — and the escrow ceremony depends on that. The hub documents the dependency in the very
panel that shows it as „waiting".
## The question this does NOT yet answer, and how it is being measured
Whether the box applies the descriptor **by itself**, and how long that takes. That decides
everything about severity: a few minutes of convergence makes the new guide step slightly too early
in the journey; never converging without an operator press makes it a broken promise on every fresh
box. A watcher is now polling the box for `pvesm` gaining a PBS storage and the agent config gaining
`escrow.pbs_storage_id`, and the ceremony will be retried when it does. Nothing was pressed, and the
box is being left to do it alone.
@@ -0,0 +1,62 @@
session=64 csrf=64
--- preflight ---
{"data":{"agent_supported":true,"escrow_state":"pending","items":[{"id":"pbs_storage_id","ok":true,"detail":"felhom-pbs"},{"id":"dr_tier","ok":true,"detail":"DR tier applied"},{"id":"age_binary","ok":true,"detail":"/usr/bin/age"},{"id":"hub_upload","ok":true,"detail":"hub upload target configured"},{"id":"staged_secret","ok":true,"detail":"staged secret present"},{"id":"sudo_grant","ok":true,"deta
--- start (re-authenticates with the dashboard password) ---
POST /api/escrow/start -> 200
{"data":{"job_id":"escrow-1789591140971464235","phase":"running"},"error":"","ok":true}
--- poll ---
{"data":{"claim_expires_in_sec":598,"claimable":true,"claimed":false,"detail":"","entropy_bits":129.24070185585344,"job_id":"escrow-1789591140971464235","key_fingerprint":"6b:ca:5f:3f:ca:0f:e2:3f:fb:2
--- claim (ONE-SHOT reveal; value goes to a file, never to stdout) ---
POST /api/escrow/claim -> 200
recovery code claimed: 83 chars, 10 words — written to a file, NOT printed
--- escrow state after the ceremony ---
## THE ANSWER: the gap is a TIMING gap, and the box closes it BY ITSELF
The watcher left the box alone and polled. Nothing was pressed.
20:30:12Z pbs_storage=none escrow.pbs_storage_id=none
20:31:13Z pbs_storage=none escrow.pbs_storage_id=none
20:32:14Z pbs_storage=none escrow.pbs_storage_id=none
20:33:15Z pbs_storage=none escrow.pbs_storage_id=none
20:34:15Z pbs_storage=none escrow.pbs_storage_id=none
**20:35:16Z pbs_storage=felhom-pbs escrow.pbs_storage_id=felhom-pbs -> CONVERGED**
That is **~17 minutes after the bind** (20:18:15Z) and ~6 minutes after the ceremony first refused.
The retry then passed every preflight item:
pbs_storage_id ok (felhom-pbs) · dr_tier ok (DR tier applied) · age_binary ok (/usr/bin/age)
· hub_upload ok · staged_secret ok · sudo_grant ok · agent_supported true · escrow_state pending
POST /api/escrow/start -> 200, job escrow-1789591140971464235, phase running
GET /api/escrow/status -> claimable:true, entropy_bits 129.2, key fingerprint 6b:ca:5f:3f:…
POST /api/escrow/claim -> **200** — the recovery code, 83 characters / 10 words, revealed ONCE
(written to a 0600 file; not printed, not committed — the household writes it on paper)
**So the product is not broken here, and the fix shipped tonight is not wrong — the GUIDE's timing
is.** `VOLUNTEER-first-hour.md` §6 (written a few hours ago) places the recovery code immediately
after the dashboard password and before the first app. On a fresh box that moment is inside the
~17-minute window where the agent has not yet applied the DR descriptor, so a volunteer following the
guide literally meets „exit 2: selftest=escrow-create requires -storage <pbs-storage-id>" and a 409,
with nothing on the page telling them to simply wait a quarter of an hour.
The escrow reminder bar (R-543, also shipped tonight) makes this sharper rather than softer: it is on
every page urging the household to do the very thing that cannot yet be done.
## The state flipped: escrow_state = **escrowed**
Read from the box's own settings after the ceremony:
enabled=True host=u629488-sub4.your-storagebox.de
**escrow_state=escrowed**
last_run=None last_status=None snapshot_count=None (never run yet — honest, not "0")
„Claimed" and „escrowed" are different facts: the claim is the household seeing the code once, the
flip to `escrowed` happens when the hub's ACK confirms `sha256(local repo_password)` matches the
stored escrow. Both happened. The off-site tier is therefore ARMED for the first time on this box,
and the escrow reminder bar shipped tonight should now be gone from every page — which the first
round will read back rather than assume.
## An alarm caused by MY recovery, flagged so it is never scored as a round's alarm
Sep 16 20:39 **error storage_disconnected** „Meghajtó váratlanul leválasztva: Adatlemez"
That is the guest restart I performed to clear the read-only wedge. It is a TRUE alarm — the drive
really did go away for those seconds — but it belongs to my repair, not to any accident in the
schedule. Round 11's drawn accident is `drive pulled 20min`, and when that round is scored this
20:39 event must not be mistaken for it.
@@ -0,0 +1,9 @@
## events baseline, captured 2026-09-16T20:15:23Z, BEFORE round 1
## everything above this line in later dumps belongs to yesterday's box, not tonight's
| Time | Severity | Type | Message | Source
| Sep 16 18:47 | error | node_down | No report received for 1h | hub
| Sep 16 18:17 | info | selfbind_link_sent | Self-bind link e-mailed (host delete) | hub
| Sep 16 18:17 | warning | node_stale | No report received for 30m | hub
| Sep 16 18:16 | warning | host_stale | Host tester-1-33b6a9: no report for 30m | hub
| Sep 16 17:25 | info | pbsdr_adopted | PBS DR token adopted by operator re-issue (endpoint held a token, hub had no descriptor) | hub
| Sep 16 17:16 | info | app_deployed | Alkalmazás telepítve: Nextcloud | controller
@@ -0,0 +1,66 @@
2026-09-16T20:11:58Z first-boot watcher started (box 192.168.0.115, VM 336)
2026-09-16T20:11:58Z PVE port 8006 answers
2026-09-16T20:14:57Z watcher restarted (the previous one killed ITSELF: pkill -f matched its own command line)
2026-09-16T20:14:58Z bootstrap=activating agent=none guests=0
2026-09-16T20:15:20Z bootstrap=activating agent=none guests=0
2026-09-16T20:15:42Z bootstrap=activating agent=none guests=0
2026-09-16T20:16:04Z bootstrap=activating agent=none guests=0
2026-09-16T20:16:25Z bootstrap=activating agent=none guests=0
2026-09-16T20:16:47Z bootstrap=activating agent=none guests=0
2026-09-16T20:17:09Z bootstrap=activating agent=none guests=0
2026-09-16T20:17:30Z bootstrap=activating agent=none guests=0
2026-09-16T20:17:52Z bootstrap=activating agent=none guests=0
2026-09-16T20:18:14Z bootstrap=activating agent=none guests=0
2026-09-16T20:18:36Z bootstrap=activating agent=none guests=0
2026-09-16T20:18:58Z bootstrap=activating agent=none guests=0
2026-09-16T20:19:27Z bootstrap= agent=none guests=0
2026-09-16T20:19:56Z bootstrap= agent=none guests=0
2026-09-16T20:20:25Z bootstrap= agent=none guests=0
2026-09-16T20:20:54Z bootstrap= agent=none guests=0
2026-09-16T20:21:23Z bootstrap= agent=none guests=0
2026-09-16T20:21:50Z bootstrap= agent=none guests=0
2026-09-16T20:22:19Z bootstrap= agent=none guests=0
2026-09-16T20:22:48Z bootstrap= agent=none guests=0
2026-09-16T20:23:17Z bootstrap= agent=none guests=0
2026-09-16T20:23:46Z bootstrap= agent=none guests=0
2026-09-16T20:24:15Z bootstrap= agent=none guests=0
2026-09-16T20:24:44Z bootstrap= agent=none guests=0
2026-09-16T20:25:12Z bootstrap= agent=none guests=0
2026-09-16T20:25:41Z bootstrap= agent=none guests=0
2026-09-16T20:26:10Z bootstrap= agent=none guests=0
2026-09-16T20:26:38Z bootstrap= agent=none guests=0
2026-09-16T20:27:07Z bootstrap= agent=none guests=0
2026-09-16T20:27:36Z bootstrap= agent=none guests=0
2026-09-16T20:28:05Z bootstrap= agent=none guests=0
2026-09-16T20:28:32Z bootstrap= agent=none guests=0
2026-09-16T20:29:01Z bootstrap= agent=none guests=0
2026-09-16T20:29:30Z bootstrap= agent=none guests=0
2026-09-16T20:29:59Z bootstrap= agent=none guests=0
2026-09-16T20:30:28Z bootstrap= agent=none guests=0
2026-09-16T20:30:58Z bootstrap= agent=none guests=0
2026-09-16T20:31:27Z bootstrap= agent=none guests=0
2026-09-16T20:31:56Z bootstrap= agent=none guests=0
2026-09-16T20:32:25Z bootstrap= agent=none guests=0
2026-09-16T20:32:54Z bootstrap= agent=none guests=0
2026-09-16T20:33:23Z bootstrap= agent=none guests=0
2026-09-16T20:33:52Z bootstrap= agent=none guests=0
2026-09-16T20:34:21Z bootstrap= agent=none guests=0
2026-09-16T20:34:50Z bootstrap= agent=none guests=0
2026-09-16T20:35:19Z bootstrap= agent=none guests=0
2026-09-16T20:35:48Z bootstrap= agent=none guests=0
2026-09-16T20:36:17Z bootstrap= agent=none guests=0
2026-09-16T20:36:46Z bootstrap= agent=none guests=0
2026-09-16T20:37:15Z bootstrap= agent=none guests=0
2026-09-16T20:37:44Z bootstrap= agent=none guests=0
2026-09-16T20:38:13Z bootstrap= agent=none guests=0
2026-09-16T20:38:42Z bootstrap= agent=none guests=0
2026-09-16T20:39:11Z bootstrap= agent=none guests=0
2026-09-16T20:39:40Z bootstrap= agent=none guests=0
2026-09-16T20:40:09Z bootstrap= agent=none guests=0
2026-09-16T20:40:38Z bootstrap= agent=none guests=0
2026-09-16T20:41:07Z bootstrap= agent=none guests=0
2026-09-16T20:41:36Z bootstrap= agent=none guests=0
2026-09-16T20:42:05Z bootstrap= agent=none guests=0
2026-09-16T20:42:25Z --- bootstrap journal, last 40 ---
2026-09-16T20:42:28Z --- what it landed on ---
2026-09-16T20:42:31Z watcher done
@@ -0,0 +1,17 @@
## 2026-09-16T19:52:02Z golden 0.245.0 — bake in the drill VM (RUNBOOK-manual-build §4.0/§4.1)
reverted to virgin (this also proves no qemu holds the qcow2)
cold boot started
ssh up: pve-manager/9.2.2/b9984c6d90a4bd80 (running kernel: 7.0.2-6-pve)
template: debian-13-standard_13.6-1_amd64.tar.zst
template downloaded
script + token landed, non-empty, executable
bake launched 2026-09-16T19:52:47Z (GITEA_USER=admin — the 'kisfenyo' namespace error cost a bake)
token leak check on the unit (needle proven non-empty, so the grep cannot match everything): 0
unit state: inactive at 2026-09-16T19:58:15Z
log copied off the machine FIRST: 325 lines
markers: overlay2=1 mountpoints=2 upload=1 FATAL=0 skipped=0
GOLDEN_VERSION=0.245.0
GOLDEN_SHA256=7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626
REGISTRY CHECK (the outcome, not the attempt): golden 0.245.0 -> http=200
PUBLISHED; guest destroyed, qemu exited, disk reverted to virgin
## bake finished 2026-09-16T19:58:52Z
@@ -0,0 +1,21 @@
install pressed at 2026-09-16T20:07:11Z
2026-09-16T20:07:33Z watcher started; completion is judged from the DISK (blocks actually used), never from the screen
2026-09-16T20:07:34Z system-disk blocks used: 3233 MiB (stable checks: 0)
2026-09-16T20:08:04Z system-disk blocks used: 5375 MiB (stable checks: 0)
2026-09-16T20:08:35Z system-disk blocks used: 6762 MiB (stable checks: 0)
2026-09-16T20:09:06Z system-disk blocks used: 6843 MiB (stable checks: 0)
2026-09-16T20:09:36Z system-disk blocks used: 6942 MiB (stable checks: 0)
2026-09-16T20:10:07Z system-disk blocks used: 6942 MiB (stable checks: 1)
2026-09-16T20:10:37Z system-disk blocks used: 6942 MiB (stable checks: 2)
2026-09-16T20:11:08Z system-disk blocks used: 6942 MiB (stable checks: 3)
2026-09-16T20:11:08Z INSTALL COMPLETE (disk stopped growing at 6942 MiB)
2026-09-16T20:11:08Z applying the post-install fix: stop, detach the CD, boot order scsi0, start from disk
disk usage before: 6942 MiB
update VM 336: -delete ide2
update VM 336: -boot order=scsi0
boot: order=scsi0
scsi0: nvme-scratch:336/vm-336-disk-1.raw,size=32G
scsi1: nvme-scratch:336/vm-336-disk-0.raw,size=100G
scsihw: virtio-scsi-single
started from disk at 2026-09-16T20:11:19Z
2026-09-16T20:11:19Z post-install fix done
@@ -0,0 +1,362 @@
## CHAOS NIGHT — Phase 0 notes (2026-09-16 evening, CEST)
### Baselines, re-verified live against Gitea at 21:49 CEST (not copied from the brief)
felhom-controller 714d5bce0920 v0.245.0 (MinAgent 0.131.0) — clean, in sync
felhom-agent e98b857684f4 v0.131.0 — clean, in sync
felhom.eu d124c77e176d hub v0.116.0, ISO 1.28.0 live — clean, in sync
app-catalog 94bc5febaca2 — clean, in sync
Register: highest R-545, 212 open. Golden waiver valid to 2026-09-27.
### A claim in the brief, CHECKED rather than inherited
The brief says the customer `tester-1` has "no host". CONFIRMED: /hosts lists exactly three hosts —
demo-felhom-8363b5, demo-hp-bb76ea, drill-r50-0a4f9a. There is no tester-1 host record. The
customers list showing „tester-1 … DOWN … 0.244.0" is the customer's LAST KNOWN state, not a live
box; reading that row as a host record would have been the mistake.
### 0.1 — golden 0.245.0, baked and published
Launched 19:52:47Z as a transient unit in the drill VM; finished 19:58:15Z.
markers: overlay2=1 mountpoints=2 upload=1 FATAL=0 publish-skipped=0
GOLDEN_VERSION=0.245.0
GOLDEN_SHA256=7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626
REGISTRY CHECK (the outcome, not the attempt): golden 0.245.0 -> http=200
then: guest 9100 destroyed, token shredded, qemu exited, disk reverted to virgin.
All three of 2026-09-16's bake failures were guarded against and none recurred: `scp -P` (not the
ssh `-p`), `chmod 0700` on the script, and `GITEA_USER=admin` (not the first credentials line).
The token-leak check ran with a needle PROVEN non-empty first, because `grep -F ""` matches every
line and an instrument that reports a hit on an empty needle is not a measurement.
HONEST NOTE ON THE EXIT CODE: the wrapper script exited **144** while every measured outcome was
good. That is why teardown was gated on the REGISTRY answering 200 and not on an exit code — this
repo's own "exit codes that lie" class. The artifact is published and verified independently.
### 0.1b — vouched in the hub
POST /configuration/artifacts -> 303, then READ BACK from the page (the outcome, not the POST code):
golden currently vouched: 0.245.0
golden sha now: 7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626
agent 0.131.0 and the wrapper sha left exactly as they were.
DECISION, and the reason, because silence reads as agreement: the GLOBAL controller floor was left at
0.244.0 and NOT raised to 0.245.0. The new box installs golden 0.245.0, which already carries
controller 0.245.0, so the floor is not needed to deliver anything tonight; raising it would have
pushed a controller update onto demo-felhom, a box that is not part of this drill, at 22:00 at night.
The brief asked for a bake and a vouch, not a floor raise.
### 0.2 — the box
demo-hp (Tier 0). VM 336 „tester1-chaos-night": 8 GiB, 4 cores, cpu host, virtio-scsi-single,
scsi0 = 32 G system disk, scsi1 = 100 G data disk, both on `nvme-scratch` (dir storage, path
/mnt/hdd_1, is_mountpoint yes — the NVMe at its ROOT, per target-selection.md). CD-ROM is the
PUBLISHED felhom-installer-1.28.0-pve9.2-1.iso. **The boot order was set in its own `qm set`** —
combining it with the disk call silently yields `order=net0;ide2` and the VM netboots.
boot: order=ide2;scsi0 (verified by reading `qm config 336` back)
The VM took DHCP 192.168.0.115 and booted the installer's graphical entry.
### Console driving — measured, not assumed
* Enter on the EULA page: advances.
* Enter on the Location page: lands IN the Country text field and does NOT press Next (the
documented GTK trap).
* **Alt+N works as the „Next" mnemonic** and is what actually drives this installer. Recorded
because the previous drill switched to the text-mode entry to avoid exactly this problem.
* The VM has a QEMU HID Tablet (absolute), so `mouse_move <x> <y>` takes absolute coordinates and
a real click is available as a fallback. `info mice` says so — checked, not assumed.
* The installer REFUSES the prefilled `mail@example.invalid` with „Please enter a valid Email
address" and simply does not advance. The dialog is above the fold, which is why the cropped
strip looked like „nothing happened" — the full screen showed the reason.
* Keyboard layout is Hungarian (QWERTZ): `y` and `z` are swapped and `@` is AltGr+V. The drill root
password was generated from a-x plus digits so it is identical under either layout, and it is
stored out-of-band (0600, scratchpad) — never in a committed file.
### Console driving on a Hungarian-layout installer — MEASURED, and it cost four round trips
These are facts about driving a PVE 9.2 graphical installer headlessly through `qm monitor`, and
every one of them was measured on this box tonight rather than recalled:
* `sendkey alt-n` is the „Next" mnemonic and is what actually advances the installer. Plain `ret`
advances ONLY the EULA page; on every later page it lands inside a text entry and does nothing.
* **`sendkey altgr-v` does NOTHING.** On a Hungarian layout `@` is AltGr+V, and the QEMU key name
that works is **`alt_r`** — `sendkey alt_r-v` types the `@`. The failure is silent: the character
is simply absent, so `admin@felhom.eu` became `adminfelhom.eu` and the installer then refused the
page with „Please enter a valid Email address" — a refusal that looks exactly like „nothing
happened" if you only crop the bottom strip of the screen.
* `mouse_move <x> <y>` did not move the pointer even though `info mice` reports a QEMU HID Tablet
(absolute) as the active device. Clicking was abandoned; keyboard navigation is the reliable path.
* **Focus is found by MEASURING it, not by counting tabs.** A 20-line script samples the blue border
of each entry box in the screendump and prints which one is focused; Tab is then pressed until
the wanted field reports focus. Counting keystrokes blind is how a password ends up in the wrong
field. (`blueness = mean(B-R)` over the box's top border: focused +58, unfocused 0.)
* The installer REFUSES the prefilled `mail@example.invalid`.
### Fence check at this point
demo-hp: /mnt/hdd_1 has 876 G free; VM 336's two raw disks are sparse (132 G apparent, 12 K actual).
`/` on demo-hp is at 77 % and was deliberately NOT used for VM disks. Guests 9201 and 9202 untouched
and running (they are the standing demo boxes, and 9202 is the scratch guest the morning restore
will use). No `local-lvm`, no prune, `drill-r50` not touched.
### A second brief claim, CHECKED rather than inherited: „the automatic mail is waiting in the mailbox"
CONFIRMED. The mailbox holds three „[Felhom] Kösd össze a Felhom dobozodat" messages to
`tester1@felhom.eu` from `monitoring@felhom.eu`, the newest at **2026-09-16T18:17:46Z** — the same
second as yesterday's host delete, which is R-509's automatic trigger firing. So the box installed
tonight should be bindable with **no operator press**, and the pre-declared intervention O1 (the
„Send self-bind link" button) should not be needed. Whether it IS needed is a measurement of this
night, not an assumption: if the waiting link turns out to be superseded or refused, that is a
finding and the press becomes O1.
The other mail in that mailbox worth noting, because it is the same customer's history and could
confuse a reader of this evidence: two „Új beállító kód — újratelepült a szervered" messages
(10:01:57Z and 16:00:58Z) and one „Beállító kód a jelszavad visszaállításához" (11:12:06Z), all from
2026-09-16 — those belong to yesterday's drills, not to tonight's box.
NOT recorded here, deliberately: the bind link itself. It carries a one-time token, and a one-time
secret does not go into a committed file (nor was the mail body fetched into the session transcript
for the same reason — the link will be taken from the hub at bind time instead).
### The one catalog bump — what it is, and why the ORDER matters
Round 7 drew `update nextcloud`, so the bump has to be on the **nextcloud** template, not on whatever
small app I would have picked. The brief asked for "one drill bump of one SMALL app"; the draw
overrides the choice of app, so the compromise is to bump the template's small sidecar pin rather
than the big application image:
templates/nextcloud/docker-compose.yml:103 redis:7-alpine -> redis:7.4-alpine
`redis:7.4-alpine` was CHECKED to exist on Docker Hub before being written down — inventing a tag
would have made round 7 fail for the wrong reason, and a round that fails for the wrong reason
measures nothing. `catalog_since` moves with it, per this repo's own rule that any commit changing an
`image:` line must.
**The bump is NOT pushed yet, and the order is the point:** nextcloud must be DEPLOYED from the
current catalog first, or the box installs 7.4-alpine immediately and round 7 has no update to
apply. Sequence: seed the twelve apps → then push the bump → the box's git-sync picks it up within
15 min (or the „Sablonok frissítése" button) → round 7 at 02:00 has a real update waiting.
Fence note: the catalog is SHARED with the demo boxes. A bump offers an update; it never applies one
— the guarded update needs a press — so the demo boxes' standing apps are not disturbed by this.
### Deploy fields, read from the templates rather than guessed
All twelve templates live under `templates/<app>/`, not at the repo root (my first lookup used the
wrong layout and returned "NO .felhom.yml" twelve times — recorded because the wrong answer was
uniform and therefore looked authoritative). Only SUBDOMAIN and HDD_PATH are marked `required`;
every secret field is optional in metadata but the server refuses a deploy without them (the
metadata-vs-server contradiction measured 2026-09-16), so the seeding script supplies generated
values for all of them and writes them to a 0600 file that is never echoed.
Apps needing the data drive: nextcloud, immich, paperless-ngx, jellyfin.
### 0.2b — the install itself
Install pressed 20:07:11Z. **Completion was judged from the DISK, not from the screen** — a PVE
install with "Automatically reboot" ticked re-enters the installer, so a finished install and a stuck
one look identical on screen. The system disk's blocks-used grew 3233 → 6942 MiB and then held
steady across three consecutive 30-second checks:
20:09:36Z 6942 MiB (stable 0) · 20:10:07Z 6942 (1) · 20:10:37Z 6942 (2) · 20:11:08Z 6942 (3)
-> INSTALL COMPLETE 20:11:08Z
Then the documented post-install fix, because a RUNNING guest keeps the boot order QEMU started
with — editing the config mid-run is not enough:
qm stop 336 · qm set 336 --delete ide2 · qm set 336 --boot order=scsi0 (its own call) · qm start
read back: `boot: order=scsi0`, no ide2, both disks intact
started from disk 20:11:19Z
The summary page was checked before Install was pressed, and the line that mattered was
**„Disk(s): /dev/sda"** — the 32 G system disk alone. The 100 G data disk was NOT offered to the
installer and was not touched. That is the check that makes pressing Install safe on a box with a
second disk, and it is the one a two-disk filter has silently got wrong elsewhere in this project.
### How this box is driven, and how the alarms are read — method facts, each one measured
* **`qm guest exec` is UNUSABLE on this box.** The VM was created with `--agent 1`, but a plain PVE
install does not run `qemu-guest-agent`, and `qm agent 336 ping` answers nothing. My first
first-boot watcher was written against `qm guest exec` and therefore sat silent while reporting
nothing — it looked like a box that would not boot, and it was an instrument pointed at nothing.
* **The access path is SSH to the box** (`root@192.168.0.115`, the drill password via `sshpass -e`
from a 0600 file, never on a command line). Confirmed with `LOGIN_OK`, hostname `chaosnight`,
`pve-manager/9.2.2`.
* **The alarm instrument is the hub's own `events-table`** on `/customers/tester-1` (note: the
customer page is `/customers/<id>`, NOT `/configs/<id>` — the latter redirects). The tabs are
client-side, so `?tab=events` returns the same document; the events must be pulled out of the
table by its container id. The hub pod has **no `sqlite3`** and `hub.db` is 298 MB with a live
`-wal`, so copying the DB is both heavy and stale-prone; the page is the correct instrument.
* **Events baseline, captured before round 1:** the newest pre-drill event is
`Sep 16 18:47 error node_down` (yesterday's box). Anything newer belongs to tonight. Without this
marker, yesterday's `node_stale`/`node_down`/`selfbind_link_sent` rows would be scored as
tonight's alarms.
### My own mistake, recorded because it cost two watchers
`pkill -f "first-boot watcher"` **matched its own command line** and killed the very background job
it was meant to clear, twice, each exiting 144. Same class as the documented `pgrep -f
qemu-system-x86_64` self-match. The replacement watcher does no pkill at all.
### Fence note on a secret
While following redirects to find the customer page, `curl -w '%{url_effective}'` printed the hub
password back to me inside the resolved URL. It is not in any file written here, and that format
option is not used again. Recorded rather than quietly dropped, because the next person will hit the
same flag.
### 0.2c — the bind, as a volunteer does it: ZERO operator presses
The box registered itself as an unclaimed appliance and then sat polling every 30 s — correctly, and
visibly: „not bound yet — polling every 30s until the operator or a customer self-bind lands". No
agent and no guest exist until the bind happens, so a reader who expected the box to finish
installing by itself would have mis-read a waiting box as a stuck one.
The volunteer's own path was taken, end to end:
* the link came from the **waiting mail** (minted 2026-09-16T18:17:46Z by yesterday's host delete,
R-509's automatic trigger) — not from the operator's „Send self-bind link" button;
* the **pairing code** was read off the box's own console: `4SY-4TX`;
* the **„Tulajdonosi jelmondat"** (5 words) came from the hub's customer record, which is where the
operator hands it from — it is never e-mailed, by design.
POST /bind/<token> -> 200 at 2026-09-16T20:18:15Z
„Sikeres összekötés. A doboz kb. egy percen belül folytatja a telepítést. Ezt az oldalt
bezárhatod — a beállítás a háttérben befejeződik, és a vezérlőpultod hamarosan elérhető lesz."
**O1 (pre-declared) was NOT used and is not counted.** The brief allowed one press of „Send self-bind
link" if no mail was waiting; a mail WAS waiting and it worked, so the bind cost zero interventions.
Secrets discipline for this step: the bind token, the passphrase and the box's root password each
live in a 0600 scratchpad file and were passed to `curl --data-urlencode name@file`, so no value
reached a command line. None of the three is written into this evidence, and the passphrase's only
description here is its shape (5 words, 35 characters).
### 0.2d — the box rotates its own root password at day-0, and my access died with it
At 20:18:58Z my SSH to the box still worked; at 20:19:04Z it answered „Permission denied", and the
background watcher lost access in the same window (its 20:19:27Z line came back empty). The bind at
20:18:15Z had started the day-0 install.
**This is the design, not a defect, and it was confirmed in source rather than guessed:**
`scripts/felhom-host-install.sh` (≈2079-2104) generates a strong `root@pam` password with `openssl
rand`, sets it through `chpasswd` on **stdin** (no argv, no log), and vaults it to the hub with
`PUT /api/v1/hosts/<id>/recovery-credential` over the enroll-authenticated channel. The password is
never logged, printed, or written to a file anywhere on the box. So the installer-time password I
typed into the Proxmox installer is dead by intent the moment day-0 runs.
**What this changes for tonight:** every accident that acts INSIDE the box (docker restart, tunnel
kill, filling the system disk) needs the hub-vaulted break-glass credential
(`/hosts/<host-id>/reveal-recovery-credential`), not the install password. Discovered at 22:19 CEST,
before the rounds began, rather than at 01:10 in the middle of round 5 — which is the only reason it
is a method note here and not an intervention later.
**How it was diagnosed honestly:** my first instinct was that I had broken my own environment. That
was ruled out first — the password file was still 21 bytes, the variable still 20 characters, and the
same credential had worked six seconds earlier. Only then was the box's own behaviour blamed, and
only after the installer source confirmed the mechanism.
### 0.2e — THE F-14 PATH, MEASURED LIVE FOR THE FIRST TIME
The brief named this as a claim that had never been measured: „the WG hook provisions by itself after
an acknowledged delete". Tonight it ran, unprompted, and the hub recorded it:
Sep 16 20:18 appliance_bound (customer_selfbind)
„Az ügyfél saját maga kötötte össze az új eszközt (bare-metal telepítés); a hozzáférést a
doboz a következő lekérdezéskor megkapja."
Sep 16 20:18 appliance_credential_delivered
„Új eszköz (bare-metal telepítés) megkapta a hozzáférést és megkezdi a beállítást."
Sep 16 20:18 claim_reissued_reenroll
„Új beállító kódot küldtünk a szerver újratelepítése után (6. generáció) az ügyfél címére."
Sep 16 20:19 offsite_reissued
„Az offsite (házon kívüli) mentési hozzáférést újra kiadtuk — az új egyszeri jelszót a vezérlő
a következő frissítéskor átveszi."
Sep 16 20:19 **pbsdr_auto_reissue**
„Previous key destroyed (acknowledged deletion) — credentials re-issued automatically."
That last line is the one that matters. Yesterday's box was removed through the **acknowledged**
delete flow, and tonight's box therefore got its off-site and PBS-DR credentials **with no operator
press at all** — exactly what the F-14 ruling of 2026-07-13 says should happen on that path, and the
half of that ruling nobody had yet watched happen.
**Consequence for this night's intervention count:** BOTH pre-declared presses are unnecessary.
O1 („Send self-bind link") was not needed because the automatic mail was waiting; O2 („Re-issue PBS
credentials") was not needed because the acknowledged-delete path re-issued by itself.
**Interventions so far: 0.**
Host enrolled as **`tester-1-022354`**, agent **0.131.0**, ONLINE, guests 0/0 at 20:20Z — the
customer guest is still being created from the golden.
### 0.2f — day-0 delivered tonight's golden, with no hand upgrade
guest: 9201 „tester-1", running, created by the bootstrap from the golden
controller image: gitea.dooplex.hu/admin/felhom-controller:**0.245.0** — „Up … (healthy)"
agent: felhom-agent **0.131.0**
host: `tester-1-022354`, ONLINE in the hub
The golden baked at 19:58Z tonight (sha 7a08aa1a…) is what this box installed. Nothing was upgraded
by hand, and the controller the customer will use is the release this drill is validating. That is
the delivery half of the chain: bake -> vouch -> a fresh box lands on it.
both disks present to the box: `sda` 32 G (system, PVE + LVM) and **`sdb` 100 G** (the data disk,
still unformatted — the household's drive, initialised through the storage page in the next step).
MY OWN MEASUREMENT ERROR, recorded: the first `lsblk` was piped through `head -12` and stopped one
line short of `sdb`. For a minute the box looked like it had NO data disk — a wrong answer produced
entirely by my own truncation, not by the box. Re-read without the pipe, `sdb 100G` is plainly there
and `qm config 336` still shows `scsi1` attached. An instrument that can cut off its own answer is
not a measurement.
### the dashboard setup code
The hub mailed a fresh „Új beállító kód — újratelepült a szervered" at **20:18:56Z** (the 6th
generation for this customer), 72 hours valid, delivered to `tester1@felhom.eu`. That is the code the
volunteer types on „A szerver beállítása" to set their own dashboard password — and it arrived by
itself, as part of the same automatic re-enrolment that needed no operator press.
### 0.2g — the dashboard claimed, by the volunteer, with the mailed code
The code from the 20:18:56Z mail („ősrégen-újraért-címbetű", 3 words, 72 h) was typed into
„A szerver beállítása" together with a 20-character password the household chooses.
POST /claim -> 302
POST /login (new pw) -> 302, and a `felhom_session` cookie was issued
**The second line is the proof; the first is only an attempt.** This repo has a standing trap that
an HTTP 200 (or a redirect) can be a refusal — the claim page re-rendering itself looks exactly like
success from the status code alone. The claim is called successful here because the password it set
then opened a session, which is the consequence a customer actually cares about.
Both values went in as FILES (`--data-urlencode name@file`, 0600, pushed with `pct push`), so
neither the setup code nor the new password ever reached a command line on the host or in the guest.
### 0.2h — tonight's release, seen working on a box that installed itself
The storage page of this fresh box carries `<form method="POST" action="/backup/escrow/banner/dismiss">`
— the **R-543 escrow reminder bar shipped in controller v0.245.0 a few hours ago**, rendering on a
box nobody had touched. It is there because the off-site tier was re-issued automatically at 20:19Z
and its escrow is not complete yet, which is precisely the state the bar exists for.
This is the first time that fix has been seen on a box that was not set up for the purpose of
testing it: the box installed itself from the published ISO, landed on tonight's golden, got its
off-site credentials with no press, and is now telling the household — on every page — that the
remote backup is paused until they create their recovery code. Creating it is the next step of the
guide, and of this drill.
### the data drive, as the box offers it
`/api/disks/candidates` reports exactly one initialisable device:
/dev/sdb — 107 374 182 400 B (100 GB), QEMU HARDDISK, data_bearing=false, mountable=false
and separately the guest's own system volume as already mounted at /mnt/sys_drive. The empty
100 GB disk is the household's drive and the only thing offered for initialisation — the
data_bearing=false flag is the guard that keeps a drive with someone's files on it out of this list.
### Schedule: Phase 0 ran long, and the rounds shift with it — recorded, not quietly re-timed
The drawn schedule puts round 1 at 23:30 CEST. Phase 0 will not be finished by then: the box was
installed, bound, claimed and landed on tonight's golden without trouble, but working out how the
storage wizard actually initialises a disk took several rounds of discovery, because the wizard
submits through JavaScript (`POST /api/storage/init`) rather than a form, and I refused to guess the
endpoint after a guessed path cost a 403 and a wrong diagnosis in an earlier drill.
**What shifts and what does not.** The SCHEDULE — which action, on which app, under which accident,
in which order — is unchanged; it was drawn from the seed before anything ran and is fixed. Only the
wall-clock start moves, and the ~25-minute spacing is kept from the new start. The brief allows a
round to wait provided the wait is recorded; this is that record. The 05:00 stop rule is unchanged,
so a late start means the night may reach fewer than twelve rounds, and the morning verdict will say
how many actually ran rather than implying all twelve did.
### 0.2i — the household's drive, initialised through the wizard's own endpoint
The wizard submits by JavaScript, not by a form: `POST /api/storage/init` with
`{device, fstype, mount_name, label, set_default, confirmed, durable_id}` and the CSRF meta token,
then polls `GET /api/storage/init/status`. Both were read off the live page and confirmed in
`internal/web/storage_handlers.go:362` before anything was sent.
POST /api/storage/init -> 200 {"phase":"formatting","started":true}
GET /api/storage/init/status -> phase **done**, started 20:27:21.79Z, updated 20:27:23.95Z
where=/mnt/felhom-drives/hdd_1, error="" reason=""
read back: /dev/sdb is ext4, durable_id `uuid:8f59ed90-c0e9-4e20-8584-d18af823c605`
`df`: /dev/sdb 98G, 2.1M used, 93G free, mounted on /mnt/felhom-drives/hdd_1
storage page: „hdd_1" labelled „Adatlemez", set as default
**The POST returning 200 is not the result** — it only says the job started. The format runs as a
background job precisely so a closed tab cannot abort it, so the phase poll is what says it worked,
and the `df` line is what says the household can use it.
**A known row met tonight, recorded rather than re-filed: R-542.** After the drive is formatted,
registered, mounted and made default, `/api/disks/candidates` STILL lists `/dev/sdb` under
`initialize` — now with `data_bearing:true, mountable:true` and its durable id. The same endpoint
also lists it under `attach`. That is exactly the behaviour R-542 describes (a registered, in-use
drive still offered under „initialize"), seen again on a fresh box.
@@ -0,0 +1,126 @@
--- the new disk as the box sees it ---
sda 32G disk
sdb 100G disk
sdc 64G disk
sda 32G disk
sdb 100G disk
sdc 64G disk
--- extend the thin pool onto it ---
pvcreate ok
vgextend ok
WARNING: Set activation/thin_pool_autoextend_threshold below 100 to trigger automatic extension of thin pools before they get full.
Logical volume pve/data successfully resized.
--- reclaim what MY failed pulls left behind (dangling layers only) ---
Total reclaimed space: 0B
--- after ---
LV LSize Data%
data <75.81g 15.57
root <13.81g
swap <3.88g
vm-9201-disk-0 32.00g 5.24
vm-9201-disk-1 70.00g 14.47
TYPE TOTAL ACTIVE SIZE RECLAIMABLE
Images 13 5 3.966GB 2.998GB (75%)
Containers 5 5 54.78kB 0B (0%)
## Recovery from MY OWN harness damage — what was changed, and why each change is a fixture change
The box was left with a 100 %-full thin pool and nine failed installs. Two fixture changes were made,
both to the drill VM, neither to the product:
1. **A third disk (64 G) was attached to the VM and the LVM thin pool extended onto it.**
before: `data` 11.80 g, Data% **100.00**
after: `data` 75.81 g, Data% **15.57**
The 32 G system disk was simply too small for a twelve-app household once thin-provisioning
over-subscribed a 32 G rootfs and a 70 G data volume onto an 11.8 G pool.
2. **The customer guest's RAM was raised 4096 -> 6144 MB.** The box has 8 GB and the guest had
half of it; the controller's memory guard counts COMMITTED memory, so twelve apps against a
3712 MB usable budget cannot fit however they are ordered.
`docker image prune -f` reclaimed **0 B** — the 3 GB `docker system df` calls "reclaimable" are
layers still referenced by the 13 pulled images, not dangling ones. Recorded because "75 %
reclaimable" reads like free space and is not.
**These are changes to the drill's own fixture, made in Phase 0 (setup), and they are not part of any
round's measurement.** The re-seed that follows installs the apps ONE AT A TIME, waiting for each to
reach `deployed: true`, which is what a household does and what the earlier parallel burst was not.
=== BEFORE: how is the guest rootfs mounted? ===
/dev/mapper/pve-vm--9201--disk--0 on / type ext4 (rw,relatime,stripe=16,emergency_ro)
/dev/mapper/pve-vm--9201--disk--1 on /var/lib/docker type ext4 (rw,relatime,stripe=16,emergency_ro)
=== restart the guest so ext4 remounts clean (the pool now has room) ===
=== AFTER: mount state ===
/dev/mapper/pve-vm--9201--disk--0 on / type ext4 (rw,relatime,stripe=16)
=== PROOF: can it actually write? (not an assumption) ===
WRITE OK
=== containers back? ===
cloudflared Up 45 seconds
felhom-controller Up 44 seconds (healthy)
filebrowser Up 43 seconds (healthy)
privatebin Up 45 seconds (healthy)
traefik Up 45 seconds
=== pool ===
data <75.81g 15.64
vm-9201-disk-0 32.00g 5.24
vm-9201-disk-1 70.00g 14.54
## A filled thin pool wedges the guest READ-ONLY, and adding space does not un-wedge it
Worth writing down beyond tonight, because the second half surprised me.
When the pool hit 100 %, both of the guest's ext4 filesystems remounted themselves with
**`emergency_ro`** — visible in the mount flags, not only in dmesg:
BEFORE:
/dev/mapper/pve-vm--9201--disk--0 on / type ext4 (rw,relatime,stripe=16,**emergency_ro**)
/dev/mapper/pve-vm--9201--disk--1 on /var/lib/docker type ext4 (rw,relatime,stripe=16,**emergency_ro**)
The symptom this produced was NOT „no space left": every deploy was refused with
„saving app config: writing /opt/docker/stacks/<app>/app.yaml.tmp: **read-only file system**"
which reads like a permissions problem and is nothing of the kind. Note the flags still say `rw` —
`emergency_ro` sits beside it, so a careless glance at `mount` says the filesystem is writable.
**Extending the thin pool from 11.8 G to 75.8 G did not clear it.** The space was there (Data% fell
to 15.6 %) and every write still failed. It took a guest restart for ext4 to mount clean:
AFTER: /dev/mapper/pve-vm--9201--disk--0 on / type ext4 (rw,relatime,stripe=16) — no emergency_ro
PROOF: `touch /opt/docker/stacks/.rwtest` -> **WRITE OK** (a real write, not an inference from flags)
all five containers back in ~45 s; pool 15.64 %
The proof line matters: „the flags look right" and „the filesystem accepts a write" are different
claims, and only the second one is the thing that was broken.
## CORRECTION — the drive gate was NOT stuck. I was reading a stale snapshot and a silent log.
For about four minutes I believed I had found a defect: the data drive was mounted (`df` showed
/dev/sdb, 98 G, on both the box and inside the guest) while the controller still recorded
`"disconnected": true, "stopped_stacks": ["immich","jellyfin","nextcloud","paperless-ngx"]`, and no
`storage_reconnected` event had appeared. I was about to file it.
**It was self-healing and it healed.** The controller's own DEBUG ring shows:
[gate] drive RETURNED /mnt/felhom-drives/hdd_1 — re-attached + restarted gate-stopped apps
Event pushed: storage_reconnected (info) — Meghajtó újra csatlakoztatva: Adatlemez
and the current state reads:
disconnected=None stopped_stacks=None
hub event, Sep 16 20:44 info storage_reconnected „Meghajtó újra csatlakoztatva: Adatlemez"
**Why I nearly got it wrong — two instrument faults at once:**
1. `driveGateLoop` runs on a **30-second ticker**; I read the settings file inside that window and
treated one sample as a settled state.
2. The gate's lines are **DEBUG**, so `docker logs` showed nothing, and I read that silence as
„the gate never ran". An absent log line is not evidence — the debug ring had the lines all
along (`/api/debug/logs?level=DEBUG`).
And a third, smaller one: my first attempt to read the ring parsed the JSON wrongly and reported
„total ring entries: 0" for a 29 503-byte response, which looked like confirmation of the silence.
Recorded in full because the wrong version of this paragraph would have been a filed P-row against a
mechanism that works.
## Tonight's own release, proven through its WHOLE lifecycle on this box (R-543)
while escrow_state=pending : the bar was on every page (measured earlier in Phase 0)
after the ceremony (escrow_state=**escrowed**), the same four pages:
/dashboard escrow-bar-hits=0
/launcher escrow-bar-hits=0
/backups/apps escrow-bar-hits=0
/storage escrow-bar-hits=0
The bar appeared while the off-site copy was paused, told the household exactly what to do, and
disappeared **for good** when they did it — on a box that installed itself from the published ISO,
with no one setting the scene for the test. „védi"/„védené" are both 0 on /backups/apps for now
because no class-A app is installed yet; that sentence is checked again once they are.
@@ -0,0 +1,147 @@
2026-09-16T20:38:31Z login ok (csrf 64)
tr: write error: Broken pipe
tr: write error: Broken pipe
2026-09-16T20:38:31Z bookstack accept=500
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/bookstack/app.yaml.tmp: open /opt/docker/stacks/bookstack/app.yaml.tmp: read-only file system
tr: write error: Broken pipe
2026-09-16T20:38:31Z gokapi accept=500
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/gokapi/app.yaml.tmp: open /opt/docker/stacks/gokapi/app.yaml.tmp: read-only file system"}
tr: write error: Broken pipe
2026-09-16T20:38:31Z homebox accept=500
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/homebox/app.yaml.tmp: open /opt/docker/stacks/homebox/app.yaml.tmp: read-only file system"}
2026-09-16T20:38:32Z uptime-kuma accept=500
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/uptime-kuma/app.yaml.tmp: open /opt/docker/stacks/uptime-kuma/app.yaml.tmp: read-only file sy
2026-09-16T20:38:32Z mealie accept=500
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/mealie/app.yaml.tmp: open /opt/docker/stacks/mealie/app.yaml.tmp: read-only file system"}
tr: write error: Broken pipe
2026-09-16T20:38:32Z vaultwarden accept=500
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/vaultwarden/app.yaml.tmp: open /opt/docker/stacks/vaultwarden/app.yaml.tmp: read-only file sy
tr: write error: Broken pipe
tr: write error: Broken pipe
2026-09-16T20:38:32Z adventurelog accept=500
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/adventurelog/app.yaml.tmp: open /opt/docker/stacks/adventurelog/app.yaml.tmp: read-only file
tr: write error: Broken pipe
tr: write error: Broken pipe
tr: write error: Broken pipe
2026-09-16T20:38:32Z nextcloud accept=500
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/nextcloud/app.yaml.tmp: open /opt/docker/stacks/nextcloud/app.yaml.tmp: read-only file system
tr: write error: Broken pipe
tr: write error: Broken pipe
tr: write error: Broken pipe
2026-09-16T20:38:32Z paperless-ngx accept=500
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/paperless-ngx/app.yaml.tmp: open /opt/docker/stacks/paperless-ngx/app.yaml.tmp: read-only fil
2026-09-16T20:38:32Z jellyfin accept=500
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/jellyfin/app.yaml.tmp: open /opt/docker/stacks/jellyfin/app.yaml.tmp: read-only file system"}
tr: write error: Broken pipe
2026-09-16T20:38:32Z immich accept=500
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/immich/app.yaml.tmp: open /opt/docker/stacks/immich/app.yaml.tmp: read-only file system"}
2026-09-16T20:38:32Z === re-seed finished; final state ===
bookstack deployed: false
gokapi deployed: false
homebox deployed: false
immich deployed: false
jellyfin deployed: false
nextcloud deployed: false
paperless-ngx deployed: false
privatebin deployed: true
uptime-kuma deployed: false
vaultwarden deployed: false
cloudflared felhom-controller filebrowser privatebin traefik
curl: option --data-urlencode: error encountered when reading a file
curl: try 'curl --help' or 'curl --manual' for more information
grep: /tmp/.l: No such file or directory
2026-09-16T20:41:05Z login ok (csrf 0)
2026-09-16T20:41:06Z bookstack accept=401
REFUSED: {"ok":false,"error":"authentication required"}
2026-09-16T20:41:06Z gokapi accept=401
REFUSED: {"ok":false,"error":"authentication required"}
2026-09-16T20:41:06Z homebox accept=401
REFUSED: {"ok":false,"error":"authentication required"}
2026-09-16T20:41:06Z uptime-kuma accept=401
REFUSED: {"ok":false,"error":"authentication required"}
2026-09-16T20:41:06Z mealie accept=401
REFUSED: {"ok":false,"error":"authentication required"}
2026-09-16T20:41:06Z vaultwarden accept=401
REFUSED: {"ok":false,"error":"authentication required"}
2026-09-16T20:41:06Z adventurelog accept=401
REFUSED: {"ok":false,"error":"authentication required"}
2026-09-16T20:41:06Z nextcloud accept=401
REFUSED: {"ok":false,"error":"authentication required"}
2026-09-16T20:41:06Z paperless-ngx accept=401
REFUSED: {"ok":false,"error":"authentication required"}
2026-09-16T20:41:06Z jellyfin accept=401
REFUSED: {"ok":false,"error":"authentication required"}
2026-09-16T20:41:06Z immich accept=401
REFUSED: {"ok":false,"error":"authentication required"}
2026-09-16T20:41:06Z === re-seed finished; final state ===
bookstack deployed: false
gokapi deployed: false
homebox deployed: false
immich deployed: false
jellyfin deployed: false
nextcloud deployed: false
paperless-ngx deployed: false
privatebin deployed: true
uptime-kuma deployed: false
vaultwarden deployed: false
cloudflared felhom-controller filebrowser privatebin traefik
2026-09-16T20:41:58Z login ok (csrf 64)
## My THIRD harness error, and the worst-shaped one
The guest restart that un-wedged the filesystem also cleared `/tmp` — where the dashboard password
file lived. The re-seed script's login therefore failed:
curl: option --data-urlencode: error encountered when reading a file
grep: /tmp/.l: No such file or directory
login ok (csrf **0**) <- it said "login ok" with an empty token
and every one of the eleven deploys came back:
accept=401 {"ok":false,"error":"authentication required"}
**Eleven lines that look exactly like the product refusing eleven installs.** They are nothing of the
kind: the product was right to refuse an unauthenticated caller, and the missing credential was mine.
Had I skimmed this output I would have written up a spectacular false finding — „the box refuses
every deploy after a restart" — and it would have been entirely an artefact of my own tooling.
Three of my errors tonight share one shape: **a script that keeps going after its own precondition
failed, and prints a confident line anyway** („all twelve deploys ACCEPTED", „login ok (csrf 0)",
and the `head -12` that hid a disk). The fix applied here is the one that should have been there from
the first line: the script now ABORTS when the session token is empty, and says the fault is mine
rather than reporting deploys as refused.
2026-09-16T20:41:58Z bookstack accept=202
2026-09-16T20:42:49Z bookstack INSTALLED (containers matching: 2)
2026-09-16T20:42:49Z gokapi accept=202
2026-09-16T20:42:59Z gokapi INSTALLED (containers matching: 1)
2026-09-16T20:42:59Z homebox accept=202
2026-09-16T20:43:09Z homebox INSTALLED (containers matching: 1)
2026-09-16T20:43:09Z uptime-kuma accept=202
2026-09-16T20:44:00Z uptime-kuma INSTALLED (containers matching: 1)
2026-09-16T20:44:00Z mealie accept=202
2026-09-16T20:44:50Z mealie INSTALLED (containers matching: 1)
2026-09-16T20:44:50Z vaultwarden accept=202
2026-09-16T20:45:01Z vaultwarden INSTALLED (containers matching: 1)
2026-09-16T20:45:01Z adventurelog accept=202
2026-09-16T20:46:01Z adventurelog INSTALLED (containers matching: 3)
2026-09-16T20:46:01Z nextcloud accept=202
2026-09-16T20:46:11Z nextcloud INSTALLED (containers matching: 3)
2026-09-16T20:46:12Z paperless-ngx accept=202
2026-09-16T20:46:32Z paperless-ngx INSTALLED (containers matching: 0)
2026-09-16T20:46:32Z jellyfin accept=202
2026-09-16T20:46:42Z jellyfin INSTALLED (containers matching: 1)
2026-09-16T20:46:42Z immich accept=202
2026-09-16T20:46:52Z immich INSTALLED (containers matching: 4)
2026-09-16T20:46:52Z === re-seed finished; final state ===
adventurelog deployed: true
bookstack deployed: true
gokapi deployed: true
homebox deployed: true
immich deployed: true
jellyfin deployed: true
mealie deployed: true
nextcloud deployed: true
paperless-ngx deployed: true
privatebin deployed: true
uptime-kuma deployed: true
vaultwarden deployed: true
adventurelog adventurelog-frontend adventurelog-postgres bookstack bookstack-db cloudflared felhom-controller filebrowser gokapi homebox immich-machine-learning immich-postgres immich-redis immich-server jellyfin mealie nextcloud nextcloud-db nextcloud-redis paperless-postgres paperless-redis paperless-webserver privatebin traefik uptime-kuma vaultwarden
@@ -0,0 +1,169 @@
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
login ok, csrf len 64
tr: write error: Broken pipe
tr: write error: Broken pipe
tr: write error: Broken pipe
tr: write error: Broken pipe
tr: write error: Broken pipe
2026-09-16T20:30:18Z nextcloud -> 202 {"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
tr: write error: Broken pipe
2026-09-16T20:30:18Z immich -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá
tr: write error: Broken pipe
tr: write error: Broken pipe
2026-09-16T20:30:18Z bookstack -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá
2026-09-16T20:30:19Z privatebin -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá
2026-09-16T20:30:19Z gokapi -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá
tr: write error: Broken pipe
2026-09-16T20:30:19Z vaultwarden -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá
tr: write error: Broken pipe
tr: write error: Broken pipe
2026-09-16T20:30:19Z paperless-ngx -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá
2026-09-16T20:30:19Z jellyfin -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá
2026-09-16T20:30:19Z mealie -> 400 {"ok":false,"error":"Nincs elég memória az alkalmazás telepítéséhez. Szükséges: 200 MB, Elérhető: 136 MB (öss
2026-09-16T20:30:19Z uptime-kuma -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá
tr: write error: Broken pipe
tr: write error: Broken pipe
2026-09-16T20:30:19Z adventurelog -> 400 {"ok":false,"error":"Nincs elég memória az alkalmazás telepítéséhez. Szükséges: 100 MB, Elérhető: 86 MB (össz
tr: write error: Broken pipe
2026-09-16T20:30:20Z homebox -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá
all twelve deploys ACCEPTED (202 = accepted, not installed — R-536: the two are different things)
perl: warning: Setting locale failed.
perl: warning: Please check that your locale settings:
LANGUAGE = (unset),
LC_ALL = (unset),
LC_CTYPE = "UTF-8",
LC_NUMERIC = (unset),
LC_COLLATE = (unset),
LC_TIME = (unset),
LC_MESSAGES = (unset),
LC_MONETARY = (unset),
LC_ADDRESS = (unset),
LC_IDENTIFICATION = (unset),
LC_MEASUREMENT = (unset),
LC_PAPER = (unset),
LC_TELEPHONE = (unset),
LC_NAME = (unset),
LANG = "en_US.UTF-8"
are supported and installed on your system.
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
## CORRECTION — the line my own script printed is FALSE, and it is corrected here before anything else
My seeding script ended with „all twelve deploys ACCEPTED". **Ten were accepted; two were refused.**
The sentence was a fixed `echo` at the end of the script, printed without consulting a single result
— the exact "a script that announces a conclusion it never checked" failure this project keeps
re-learning, and it would have put a false line into tonight's record.
ACCEPTED (202), 10: nextcloud · immich · bookstack · privatebin · gokapi · vaultwarden ·
paperless-ngx · jellyfin · uptime-kuma · homebox
REFUSED (400), 2:
mealie „Nincs elég memória az alkalmazás telepítéséhez. Szükséges: 200 MB, Elérhető: 136 MB…"
adventurelog „Nincs elég memória az alkalmazás telepítéséhez. Szükséges: 100 MB, Elérhető: 86 MB…"
Note also that nine of the ten acceptances carried a warning of their own:
„Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát…"
so the box was telling the truth about its memory on the way up, and then refused the last two
outright. The refusal is the controller's memory guard doing its job, fail-closed and in plain
Hungarian with the numbers in it.
**Consequence for the drawn schedule, stated now rather than discovered at 23:30:** rounds 1 and 4
act on `adventurelog` and `mealie`, and round 6 on `adventurelog` — apps that are NOT installed.
The schedule was drawn before the night and is not being re-drawn; what changes is that those rounds
must either act on an app that exists, or be recorded as unrunnable. That decision is made and
recorded explicitly, not silently.
Also true and worth keeping: „202 = accepted, not installed" (R-536). How many of the ten actually
finished installing is a separate measurement, taken next.
## The memory facts behind the two refusals — measured
the VM (the whole box): 7939 MB total, 5814 MB available
the CUSTOMER GUEST (LXC 9201): **memory: 4096, swap: 512, cores: 3**
inside the guest at the moment of the refusals: 4096 MB total, ~3685 MB available
containers actually running then: 4 (cloudflared, felhom-controller, filebrowser, traefik)
So the box had ~5.8 GB free while the guest the apps live in was capped at 4 GB — and the guard
refused the eleventh and twelfth apps against the guest's cap, counting the memory already COMMITTED
by ten in-flight installs rather than the memory currently in use. That is the honest way to count
it (otherwise ten simultaneous pulls would all be admitted and then fight), and the message quoted
the two numbers it compared.
**This is the constraint that decides whether tonight's household can be twelve apps at all.** The
brief asks for the twelve of BIGNIGHT. The box as installed gives its customer guest 4 GB. Nothing
has been changed yet: first the ten in-flight installs are allowed to finish, because the memory
picture during a pull is not the memory picture afterwards, and a decision taken on the wrong
picture is worse than a late one.
## DECISION — what happens to the rounds that name apps which are not installed
Three rounds name apps the box refused to install: round 1 (`offsite-run adventurelog`), round 4
(`offsite-run mealie`), round 6 (`backup-system adventurelog`).
**The schedule is not re-drawn.** It was fixed from the seed before anything ran, and re-drawing it
now — after seeing which apps happened to fit in memory — is exactly the "choose the night after the
fact" failure the seed exists to prevent.
What the rounds actually do, and why this costs less than it looks:
* **`offsite-run` is a TIER action, not an app action.** The off-site leg is repo-global — one
`LastRun` for the whole repository, no per-app run time — so rounds 1 and 4 exercise the tier
exactly as drawn. The named app is which app's row I read afterwards; where that app is absent,
the round records the tier's own result and says the app was not installed.
* **`backup-system` (round 6) is a whole-box action** and does not depend on the named app either.
So all three rounds run as drawn; what changes is that their "what the customer saw" cell reports the
tier or the whole-system page rather than that app's row. Each affected round says so in its own line
rather than leaving a reader to assume the app was there.
**And the memory refusal is itself a finding, not just an inconvenience:** a fresh box built to the
documented shape gives its customer guest 4 GB, and the twelve-app household of BIGNIGHT does not fit
in it. Nothing was resized to make the drill comfortable.
## Deploy progress, and a second instrument lesson
At 20:32:59Z, twelve minutes after the ten deploys were accepted:
app.yaml recorded: 10 (the box registered all ten)
containers running: 5 (cloudflared, felhom-controller, filebrowser, traefik, **privatebin**)
guest memory available: 3818 MB
So exactly one of the ten household apps was actually up; the rest were still pulling images. This is
R-536's distinction in the flesh: ten "telepítés elindítva" acceptances, one installed app.
**The instrument lesson (my second tonight):** I read the deploy watcher with
`tail -6 … | grep -vE "locale|…"`, and the locale warnings filled the whole tail, so the watcher's
ONE real line was filtered out and the watcher looked dead. I had already declared one watcher dead
tonight for a different reason and killed it with a `pkill` that killed itself. The fix both times
was the same: **ask the box directly instead of trusting my own reporting layer.** The box answered
in one call, and the watcher turned out to have been working the whole time.
@@ -0,0 +1,17 @@
#!/bin/bash
# Run on demo-hp AFTER the PVE install finishes.
#
# Why this exists: "Automatically reboot after successful installation" is ticked and the boot order
# is ide2;scsi0, so the box reboots straight back INTO the installer — and a completed install then
# looks exactly like a stuck one. A running guest also keeps the QEMU boot order it started with, so
# editing the config mid-run is not enough: the VM must be stopped.
#
# Completion is judged from the DISK, not from the screen.
set -u
V=336
echo "disk usage before: $(du -sh --block-size=1M /mnt/hdd_1/images/$V/vm-336-disk-1.raw 2>/dev/null | cut -f1) MiB"
qm stop $V; sleep 5
qm set $V --delete ide2
qm set $V --boot order=scsi0 # its OWN call, always
qm config $V | grep -E '^(boot|ide2|scsi)'
qm start $V && echo "started from disk at $(date -u +%FT%TZ)"
@@ -0,0 +1,36 @@
#!/bin/bash
# round.sh — run ONE chaos round and record the same five things.
# Usage: round.sh <n> <action> <app> <accident>
set -u
N="$1"; X="$2"; Y="$3"; Z="$4"
E=/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17
OUT="$E/round-${N}.txt"
SC=/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/adc14dfe-fc6c-4014-9378-d580d29d3595/scratchpad/chaos
say(){ echo "$(date -u +%FT%TZ) $*" | tee -a "$OUT"; }
export SSHPASS=$(cat $SC/boxroot.pw)
SSHO="-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=8 -o NumberOfPasswordPrompts=1"
BOX(){ sshpass -e ssh $SSHO root@192.168.0.115 "$@" 2>/dev/null | grep -v "Warning: Permanently"; }
say "================ ROUND $N : $X on $Y, while: $Z ================"
say "--- BEFORE: is the box steady? (every app up, hub ONLINE, no active alarm) ---"
BOX 'pct exec 9201 -- docker ps --format "{{.Names}} {{.Status}}"' | sed 's/^/ /' | tee -a "$OUT"
say " events before this round:"
bash "$E/events.sh" 3 | tee -a "$OUT"
say "--- ACTION: $X on $Y ---"
# the action itself is driven per-round by the caller's follow-up; this records the start moment
say " action start marker"
say "--- ACCIDENT: $Z (injected 10-60 s after the action starts) ---"
if [ "$Z" != "nothing" ]; then
bash "$E/inject.sh" "$Z" "$N" | sed 's/^/ /' | tee -a "$OUT"
else
say " control round — no accident, deliberately"
fi
say "--- AFTER: what the box did by itself ---"
BOX 'pct exec 9201 -- docker ps --format "{{.Names}} {{.Status}}"' | sed 's/^/ /' | tee -a "$OUT"
BOX 'pvesm status; df -h /mnt/felhom-drives/hdd_1 2>/dev/null | tail -1' | sed 's/^/ /' | tee -a "$OUT"
say "--- alarms in this round's window ---"
bash "$E/events.sh" 8 | tee -a "$OUT"
say "================ END ROUND $N ================"
@@ -0,0 +1,37 @@
# The per-round record — CHAOS NIGHT
Every round records the SAME five things, in this order, and nothing is written from memory:
1. **what the customer saw** — screens quoted verbatim (Hungarian in „ ", searched with ASCII
fragments, with a positive and a negative control; accented `grep` dies with a complexity error
here, so searching is done in Python)
2. **what the box did by itself** — no shell, no help; anything I had to do is an intervention
3. **time to steady** — measured from the accident to the moment every app is running, the hub says
ONLINE and no alarm is active. **`—` when the box never got there on its own**, never a guess
4. **which alarm fired, and was it TRUE**
5. **which alarm SHOULD have fired (per 08-alarm-ladder.md) and did not**
plus **the background loop's failures inside the round's window**, counted from its own log.
## Two things known BEFORE the night that change how rounds are scored
- **Rounds 7, 8 and 9 all block the box's network.** An event generated while the hub is unreachable
is retried 3 times over ~6 s and then **dropped permanently** (`PushEvent`, no queue). So a missing
alarm in those rounds is not evidence the alarm did not fire — it may have been posted into a
blocked path. Scored as `LOST-IN-BLOCK`, never as `MISSED`.
- **The dedupe is not one window.** 5 minutes applies ONLY to the node/host liveness events; the
default operator cooldown is **1 hour**, keyed `customer:type[...]`. A second identical alarm
inside an hour is suppressed and written to `notification_log` with status `suppressed` — so
"suppressed" and "never fired" are distinguishable, and must be distinguished.
## Steady-state check (the same command set every round)
* every app: `docker ps` shows it running AND its front door answers
* the hub: the host row reads ONLINE, guests 1/1, agent version present
* no active alarm on the dashboard; `notification_log` read for the round's window
* the data drive: the storage page reads „Aktív", not „Leválasztva"
## Evidence discipline
Evidence is copied off the box **at the end of each round, before the next accident** (R-320) — the
one that gets forgotten is the middle one, never the last.
@@ -0,0 +1,23 @@
seed: 20260917
script sha256: 4b98afe65d042df7e7dc417553b33565cfbb4afd451c68456a7cabec0858d2a1
| # | time | X — the action | Y — the app | Z — the accident |
|---|---|---|---|---|
| 1 | 23:30 | offsite-run | adventurelog | nothing |
| 2 | 23:55 | restore | gokapi | power cut |
| 3 | 00:20 | use | bookstack | disk 95% full |
| 4 | 00:45 | offsite-run | mealie | tunnel down 10min |
| 5 | 01:10 | use | privatebin | docker restarted |
| 6 | 01:35 | backup-system | adventurelog | nothing |
| 7 | 02:00 | update | nextcloud | internet gone 10min |
| 8 | 02:25 | backup-app | nextcloud | internet gone 10min |
| 9 | 02:50 | use | uptime-kuma | internet gone 10min |
| 10 | 03:15 | restore | uptime-kuma | hard reset |
| 11 | 03:40 | use | paperless-ngx | drive pulled 20min |
| 12 | 04:05 | use | paperless-ngx | nothing |
re-draw log (4 entries):
r02 X=reinstall re-drawn (nothing has been removed yet)
r06 Z=disk 95% full re-drawn (constraint 4: at most once)
r07 Z=nothing re-drawn (constraint 6: never two in a row after r2)
r08 X=reinstall re-drawn (nothing has been removed yet)
Binary file not shown.

After

Width:  |  Height:  |  Size: 25 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 157 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 137 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 131 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 131 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 10 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 134 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 11 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 136 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 11 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 137 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 150 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 133 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 43 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 9.7 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 11 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 22 KiB

@@ -0,0 +1,47 @@
#!/bin/bash
# seed_apps.sh — CHAOS NIGHT Phase 0.3: move the household in.
# Twelve apps, deployed through the controller's own API (the endpoint the UI invokes), each with
# exactly the fields its template declares. Secrets are GENERATED here and written to a 0600 file;
# none is ever echoed. Run INSIDE the guest (the controller answers on the container address).
#
# Usage: seed_apps.sh <controller-container-ip> <host-header> <pw-file> <hdd-path> <secrets-out>
set -u
B="http://${1:?ctrl ip}:8080"; H="Host: ${2:?host header}"; PWF="${3:?pw file}"
HDD="${4:?hdd path}"; SEC="${5:?secrets out}"
: > "$SEC"; chmod 600 "$SEC"
gen(){ tr -dc 'a-zA-Z0-9' </dev/urandom | head -c "${1:-24}"; }
curl -s -o /dev/null -D /tmp/.l -H "$H" -X POST --data-urlencode "password@$PWF" $B/login
SESS=$(grep -i '^set-cookie: felhom_session=' /tmp/.l | sed 's/.*felhom_session=\([^;]*\).*/\1/')
[ -n "$SESS" ] || { echo "LOGIN FAILED"; exit 1; }
C="Cookie: felhom_session=$SESS"
TOK=$(curl -s -H "$H" -H "$C" $B/backups/remote | grep -o 'name="csrf-token" content="[^"]*"' | sed 's/.*content="\([^"]*\)".*/\1/')
echo "login ok, csrf len ${#TOK}"
dep(){ # dep <stack> <json-values>
local app="$1" vals="$2"
local code
code=$(curl -s -o /tmp/.d -w '%{http_code}' -H "$H" -H "$C" -H "X-CSRF-Token: $TOK" \
-H 'Content-Type: application/json' -X POST -d "{\"values\":$vals}" \
"$B/api/stacks/$app/deploy")
printf '%s %-14s -> %s %s\n' "$(date -u +%FT%TZ)" "$app" "$code" "$(head -c 120 /tmp/.d)"
}
D='"DOMAIN":"enkicsifelhom.hu"'
NC_ADMIN=$(gen 20); PL_ADMIN=$(gen 20); GK=$(gen 20)
{ echo "nextcloud_admin=$NC_ADMIN"; echo "paperless_admin=$PL_ADMIN"; echo "gokapi_admin=$GK"; } >> "$SEC"
dep nextcloud "{$D,\"SUBDOMAIN\":\"cloud\",\"DB_PASSWORD\":\"$(gen)\",\"MYSQL_ROOT_PASSWORD\":\"$(gen)\",\"NEXTCLOUD_ADMIN_USER\":\"admin\",\"NEXTCLOUD_ADMIN_PASSWORD\":\"$NC_ADMIN\",\"HDD_PATH\":\"$HDD\"}"
dep immich "{$D,\"SUBDOMAIN\":\"photos\",\"DB_PASSWORD\":\"$(gen)\",\"HDD_PATH\":\"$HDD\"}"
dep bookstack "{$D,\"SUBDOMAIN\":\"wiki\",\"APP_KEY\":\"base64:$(gen 32)\",\"DB_PASSWORD\":\"$(gen)\"}"
dep privatebin "{$D,\"SUBDOMAIN\":\"paste\"}"
dep gokapi "{$D,\"SUBDOMAIN\":\"share\",\"GOKAPI_PASSWORD\":\"$GK\"}"
dep vaultwarden "{$D,\"SUBDOMAIN\":\"vault\",\"ADMIN_TOKEN\":\"$(gen 32)\",\"SIGNUPS_ALLOWED\":\"false\"}"
dep paperless-ngx "{$D,\"SUBDOMAIN\":\"paperless\",\"DB_PASSWORD\":\"$(gen)\",\"PAPERLESS_SECRET_KEY\":\"$(gen 32)\",\"PAPERLESS_ADMIN_USER\":\"admin\",\"PAPERLESS_ADMIN_PASSWORD\":\"$PL_ADMIN\",\"HDD_PATH\":\"$HDD\",\"PAPERLESS_OCR_LANGUAGE\":\"hun+eng\"}"
dep jellyfin "{$D,\"SUBDOMAIN\":\"media\",\"HDD_PATH\":\"$HDD\"}"
dep mealie "{$D,\"SUBDOMAIN\":\"recipes\"}"
dep uptime-kuma "{$D,\"SUBDOMAIN\":\"status\"}"
dep adventurelog "{$D,\"SUBDOMAIN\":\"travel\",\"SECRET_KEY\":\"$(gen 32)\",\"DB_PASSWORD\":\"$(gen)\"}"
dep homebox "{$D,\"SUBDOMAIN\":\"inventory\",\"HBOX_AUTH_API_KEY_PEPPER\":\"$(gen 32)\"}"
rm -f /tmp/.l /tmp/.d
echo "all twelve deploys ACCEPTED (202 = accepted, not installed — R-536: the two are different things)"
@@ -0,0 +1,45 @@
#!/bin/bash
# steady.sh — is the box back on its own? Run between every round.
#
# "Steady" is not a word here, it is a measurement: every app running, the hub says ONLINE, the data
# drive reads Aktiv, and no alarm is active. If the box did not get there BY ITSELF, the round's
# steady-state cell is `—`, never a guess and never a number I helped it reach.
#
# Usage: steady.sh <round-label> (reads the box address from $BOXIP, guest id from $GUEST)
set -u
R="${1:?round label}"
VM=336
E=/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17
OUT="$E/round-${R}-steady.txt"
H(){ ssh hp "$@" 2>/dev/null | grep -vE "locale|LC_|LANG|perl:|supported and installed|are supported"; }
say(){ echo "$(date -u +%FT%TZ) $*" | tee -a "$OUT"; }
say "=== steady check, round $R ==="
say "--- the box itself (is the VM even up?) ---"
H "qm status $VM" | tee -a "$OUT"
say "--- the customer guest and its apps (asked of the box, not of me) ---"
H "qm guest exec $VM -- bash -c 'pct list; echo ---; pct exec 9000 -- docker ps --format \"{{.Names}} {{.Status}}\"'" 2>/dev/null | tee -a "$OUT"
say "--- the hub's view: host row, guests, agent, last report ---"
cd /mnt/5_hdd/felhom.eu/git/felhom.eu
python3 scripts/read_credential.py HUB_PW /tmp/.shp >/dev/null 2>&1
HUB_PW=$(cat /tmp/.shp); IP=$(sudo kubectl -n felhom-system get svc hub -o jsonpath='{.spec.clusterIP}')
curl -s -u ":$HUB_PW" "http://$IP:8080/hosts" | python3 -c "
import sys,re,html
t=sys.stdin.read()
txt=html.unescape(re.sub(r'\s+',' ',re.sub(r'<[^>]+>',' ',re.sub(r'<script.*?</script>','',t,flags=re.S))))
i=txt.find('chaosnight')
print(' ', txt[max(0,i-40):i+220] if i>=0 else 'chaosnight NOT in the hosts table')
" | tee -a "$OUT"
say "--- alarms/events in the last 30 minutes (fired AND suppressed are different things) ---"
curl -s -u ":$HUB_PW" "http://$IP:8080/" | python3 -c "
import sys,re,html
t=sys.stdin.read()
txt=html.unescape(re.sub(r'\s+',' ',re.sub(r'<[^>]+>',' ',re.sub(r'<script.*?</script>','',t,flags=re.S))))
m=re.search(r'(Recent events|Events|Legutobbi).{0,900}', txt)
print(' ', (m.group(0)[:900] if m else txt[:400]))
" | tee -a "$OUT"
rm -f /tmp/.shp
say "=== end steady check, round $R ==="
+1
View File
@@ -728,6 +728,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-543** | **[P1-HIGH] Off-site ON by default is not off-site WORKING: on a fresh box tier 3 sits at „Kulcsletétre vár" until the household does the escrow ceremony, and nothing asks them to — while the tier-1 row now tells them their files are protected by that very copy.** MEASURED 2026-09-16 on the fresh box (controller 0.244.0, hub 0.116.0, off-site provisioned automatically by the new default): the app-backup page reads „3. mentés — Kulcsletétre vár · A távoli mentés a titkosítási kulcs letétbe helyezéséig szünetel", the remote page reads „Helyreállítási kód szükséges", and `POST /backup/offbox/run` returns 302 while producing no snapshot (the controller log shows only `offsite-credential-retry`, no restic activity). **Why it matters more than before today:** hub v0.116.0 makes off-site the default *because* a one-drive box otherwise keeps the household's files in no tier at all (R-537/R-538), and controller v0.244.0 now prints „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi" under the tier-1 row. On day one both are true-in-intent and false-in-fact: the copy is paused. **Fix shape (one of):** prompt the escrow ceremony as part of first-run when off-site is enabled and un-escrowed; and/or make the tier-1 sentence state the tier's actual state („…védené — a távoli mentés a helyreállítási kód létrehozásáig szünetel"). The ceremony itself works and is customer-facing („Helyreállítási kód létrehozása"); what is missing is that anyone is told to do it. **CLOSED 2026-09-16 — controller v0.245.0, both halves proven live.** The pause is untouched: it is the zero-knowledge escrow design, and this row was never about the mechanism. (a) **The household is asked:** while the off-site tier is configured and its escrow is not complete, every authenticated page carries „A távoli mentés szünetel, amíg nem hozod létre a helyreállítási kódot." linking `/backup/escrow`. It is the R-241 bar, second instance — same session-cookie dismissal, back at the next visit, gone for good when escrowed; **no second banner system**. It hangs off `executeTemplate`, the single render choke point, so it cannot reach only the pages someone remembered. (b) **The tier-1 sentence renders by state:** `driveFilesNoteFor` takes `tier3State`'s own vocabulary — `active` → „védi", `escrow_pending` → „védené … a helyreállítási kód létrehozásáig szünetel" + the route, no off-site and no second drive → „nincs másolat" + both ways out. **Measured live on 0.245.0:** on a paused box (9202, off-site configured through the product's own endpoint, `escrow_state=pending`) the bar renders on /dashboard, /launcher, /backups/apps and /settings; a manual `POST /backup/offbox/run` is refused by the fork-4 gate („A távoli mentés a kulcs letétbe helyezésére vár.") with `last_run=None, snapshot_count=None`; and a throwaway class-A app's row reads „…védené … szünetel" with „védi"=0. On an escrowed box (9201) the bar is absent on all three pages and the row reads „védi". Both red-proofed (the bar test fails on BOTH pages with the one hook line removed; the sentence test quotes the exact v0.244.0 promise when the state is ignored). The first-hour guide now asks for the code right after the dashboard password and before the first app. Evidence: `audits/evidence-recovery-code-2026-09-16/`. | **CLOSED 2026-09-16 — shipped in controller v0.245.0 and proven live** |
| **R-544** | **[P3-LOW] The host-delete log line says „escrow deleted: true" while the documented (and actual) effect is DEMOTION to retained custody.** MEASURED 2026-09-16 during the teardown of the fresh box: an unacknowledged delete was correctly refused 409 („has key escrow (acknowledgement missing)") and the refusal text promises the acknowledgement „moves it to retained custody"; the acknowledged delete then logged `host deleted: tester-1-33b6a9 (escrow deleted: true)`. The hub's own customer page states the truth — „host deletion only demotes custody, never destroys it… recovery-key custody is demoted to retained custody, not destroyed", with the customer delete named as „the one true purge point". **Nothing is broken; the log is.** An operator reading that line during an incident would believe a household's last key had just been destroyed, and the R-304 retention exists precisely so it is not. **Fix shape:** log what happened — `escrow custody demoted to retained (host delete)` — and keep the boolean's name out of operator-facing text. | **READY — rank P3-LOW; owner: CC (hub)** |
| **R-545** | **[P3-LOW] There is no product action that un-configures an off-site target — only one that re-starts an ORPHANED repository.** FOUND 2026-09-16 while exercising R-543's paused state on the scratch guest: `POST /backup/offbox/config` configures a target and can disable it (`enabled` unchecked), but nothing removes it. `POST /backup/offbox/reset` refuses unless `OffboxOrphaned()` is true („Az offsite tároló nincs elárvult állapotban."), and it means „start a new remote backup, set the old history aside" — not „forget this destination". So a household that sets up the wrong NAS, or a box being handed to someone else, keeps the host, user, path, ssh key and minted repo password on disk with no route to clear them; a disabled target still holds its secrets in `data/offbox/`. **Why P3 and not higher:** a disabled target runs nothing and the secrets are 0600 on the box's own disk, so nothing leaks and no copy is lost. **Fix shape:** a „Távoli cél törlése" action beside the config form that clears the target and shreds `data/offbox/`, REFUSING while the hub holds a sealed package for this box (the R-241 rule — dropping the key would orphan the history that package protects). Teardown for this session's proof had to clear it out-of-band for exactly this reason, which is the measurement. | **READY — rank P3-LOW; owner: CC** |
| **R-546** | **[P2-MEDIUM] The first-hour guide sends the household to create their recovery code at a moment when the box cannot yet do it — and the new reminder bar urges them there on every page.** MEASURED 2026-09-16/17 on a fresh box (`tester-1-022354`, guest 9201, controller 0.245.0, agent 0.131.0, installed from the published ISO 1.28.0). `VOLUNTEER-first-hour.md` §6 — added hours earlier in controller v0.245.0 — places „A helyreállítási kód" immediately after the dashboard password and **before the first app**, because until it is done the off-site copy does not run. At exactly that point the ceremony FAILS: `POST /api/escrow/start` → 200, then `GET /api/escrow/status` → `detail: "exit 2: … selftest=escrow-create requires -storage <pbs-storage-id> (or escrow.pbs_storage…"`, and `POST /api/escrow/claim` → **409** „A folyamat jelenlegi állapotában a kód nem kérhető le." **Cause, measured on both sides:** the hub had auto-provisioned the DR descriptor at 20:19 (no press — see R-534/R-511's acknowledged-delete path) and its Backup & DR panel itself read „descriptor provisioned … **waiting** · ceremony possible once the descriptor is applied on the box"; the box had no PBS storage (`pvesm status` = local + local-lvm only) and `/etc/felhom-agent/agent.json` had **no `escrow` section at all**. **It is a TIMING gap and it self-heals:** a watcher left the box alone and polled — `pbs_storage` and `escrow.pbs_storage_id` both became `felhom-pbs` at **20:35:16Z, ~17 minutes after the bind**; the retried ceremony then passed every preflight item and the claim returned 200 (83-character code, entropy 129.2 bits), and `escrow_state` flipped to `escrowed`. **Why it still matters:** for those ~17 minutes the R-543 reminder bar (also v0.245.0) is on *every* page telling the household to do the one thing that refuses, and nothing on the page says „wait a few minutes" — the volunteer meets a stderr fragment about a `-storage` flag. **Fix shape (one of):** have the escrow page/bar consult `preflight` and say „a doboz még készül — pár perc múlva próbáld újra" while `pbs_storage_id` is unset; or move the guide's step to after the first app; or make the bar appear only once preflight is green. **No product code was changed tonight** (validation run). Evidence: `audits/evidence-chaos-night-2026-09-17/phase0-escrow-failure.txt` and `phase0-escrow-retry.txt`. | **READY — rank P2-MEDIUM; owner: CC (controller copy + guide timing)** |
| **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured, and a class-A app carries one sentence saying where its files ARE protected. Proven on demo-hp through the page the customer opens: Paperless-ngx reads „1. mentés … DB + Konfig" with „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi …", while its „2. mentés" row still reads „DB + Konfig + Adatok". Red-proof: restoring the old app-shaped label fails `TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold`. **RE-PROVEN 2026-09-16 on a FRESH box** (installed from the built ISO 1.28.0, controller 0.244.0, off-site on by default): the Nextcloud row read „1. mentés … DB + Konfig" with the new sentence, „2. mentés … Nincs 2. (off-drive) másolat", „3. mentés Sikeres restic → …your-storagebox.de"; „DB + Konfig + Adatok" appeared ZERO times while the local unit held no file leg. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** |
| **R-538** | **[P1-HIGH] A tier-1 app restore reports plain success and leaves Nextcloud listing files whose bytes were never in the backup — and it destroys the app's own trash, the customer's last copy.** MEASURED 2026-09-16 on the drill box, F10 („a child deletes the photo folder"): the five photos were deleted through Nextcloud (DELETE 204, PROPFIND 404), then restored through the page exactly as a customer would (`POST /backup/restore` `stack_name=nextcloud` `snapshot_id=helyi` → 302, finished in **35 s**, „A(z) nextcloud: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult."). Afterwards the folder is back and **lists all five photos**, and **none of them opens**: `GET nyaralas-1..5` = 404 / 503×4 with `Sabre\DAV\Exception\NotFound`, while the positive controls at the same moment pass (`status.php` 200, WebDAV PUT 201, GET 200). Cause: the replayed MariaDB dump (11:01:45Z) knows the photos, the bytes live on `mp8` and were never captured (R-537). **Worse:** the bytes were still on the drive in Nextcloud's own trash (`appdata/nextcloud/admin/files_trashbin/files/Fotok.d1789556707/nyaralas-1..5.jpg`, all five present) and the restored database no longer references them — the trash listing comes back **empty**, so „restore from trash", the one route that would have worked, is gone. The customer is left with five unopenable photos, a success message, and no warning. **Fix shape:** before replaying a database whose app has an uncaptured file leg, refuse or warn („ennek az alkalmazásnak a fájljai nincsenek ebben a mentésben — a visszaállítás után a fájlok hiányozni fognak"); and never present a DB-only restore of a class-A app as a complete one. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** A unit restore refuses before anything is touched when the unit cannot return the app's drive-side files, and names the route that can. Fired live on demo-hp: `POST /backup/restore` for paperless-ngx → 302 with „Ez a mentés nem tartalmazza az alkalmazás fájljait, ezért nem állítjuk vissza az adatbázist föléjük — a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza …", and the app read `running` before AND after, so nothing was stopped and no trash was made unreachable. The database-and-settings-only path exists as a separately worded second step. Red-proof: disabling the guard fails `TestUnitRestore_RefusesWhenTheUnitCannotHoldTheFiles`. **RE-PROVEN 2026-09-16 on a FRESH box, and this time the refusal had somewhere to point:** after five photos were deleted, `POST /backup/restore` was refused with „…a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza: … „Teljes visszaállítás (fájlok + adatbázis)"", the app read `running` before AND after, and the wastebasket was untouched. The off-site route then returned all five photos — 200 with the exact uploaded sizes and sha256 IDENTICAL to the originals, 5/5, with a negative control. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseE-photos.txt`. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** |
| **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** |