CHAOS NIGHT phase 0: golden 0.245.0, a self-installing box, and R-546
gates / gates (push) Successful in 23s
The schedule was drawn from seed 20260917 and written into the findings document BEFORE round 1, with its re-draw log. Phase 0 measured: - golden 0.245.0 baked, published (registry 200, not an exit code) and vouched; the box installed itself from the published ISO 1.28.0 and landed on it with no hand upgrade (controller 0.245.0, agent 0.131.0). - ZERO operator presses: the waiting self-bind mail worked, and the acknowledged -delete path re-issued off-site AND PBS-DR credentials by itself (pbsdr_auto_reissue) - the F-14 half nobody had watched happen live. - R-546 filed (P2): tonight's own guide sends the household to create the recovery code ~17 minutes before the box can do it. It self-heals; the bar urges them there the whole time. Measured on both sides, not inferred. - R-543 proven through its whole lifecycle on a fresh box: bar present while paused, gone for good once escrowed. - Known rows met and recorded, not re-filed: R-542, R-536's failure events. Also recorded honestly: three harness errors of mine (a script that announced "all twelve deploys ACCEPTED" without checking, a "login ok (csrf 0)" that turned eleven of my own 401s into what looked like product refusals, and a head -12 that hid a disk), and a near-miss where I almost filed a defect against a drive gate that was working and logging at DEBUG. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
@@ -0,0 +1,159 @@
|
||||
# DRILL — CHAOS NIGHT: random actions on random apps while random things go wrong (2026-09-16/17)
|
||||
|
||||
**Interventions: PENDING — the run is in progress.**
|
||||
**Ready for a volunteer: PENDING.**
|
||||
**The accident-plus-action pair that hurt most: PENDING.**
|
||||
|
||||
> **Baselines, verified live against Gitea at 21:49 CEST 2026-09-16 (not copied from the brief):**
|
||||
> felhom-controller `714d5bce0920` v0.245.0 (MinAgent 0.131.0) · felhom-agent `e98b857684f4` v0.131.0 ·
|
||||
> felhom.eu `d124c77e176d` hub v0.116.0, ISO **1.28.0 published** · app-catalog `94bc5febaca2`.
|
||||
> All four trees clean and in sync. Highest register row **R-545**, 212 open. Golden waiver valid to
|
||||
> 2026-09-27. Customer **`tester-1`** (`enkicsifelhom.hu`, `tester1@felhom.eu`, no host).
|
||||
> Venue: `demo-hp` (Tier 0), a fresh nested VM, disk on the NVMe at its root. Evidence:
|
||||
> `evidence-chaos-night-2026-09-17/`.
|
||||
|
||||
## The schedule — drawn ONCE, before round 1, and written here first
|
||||
|
||||
The point of this section's position in the document is that the night could not be chosen after the
|
||||
fact. `chaos_schedule.py` is committed beside the evidence; re-running it reproduces this table.
|
||||
|
||||
- **seed:** `20260917` (the date)
|
||||
- **script sha256:** `4b98afe65d042df7e7dc417553b33565cfbb4afd451c68456a7cabec0858d2a1`
|
||||
- **generator:** `evidence-chaos-night-2026-09-17/chaos_schedule.py`, stdlib `random` seeded with the seed
|
||||
|
||||
| # | time | X — the action | Y — the app | Z — the accident |
|
||||
|---|---|---|---|---|
|
||||
| 1 | 23:30 | offsite-run | adventurelog | nothing |
|
||||
| 2 | 23:55 | restore | gokapi | power cut |
|
||||
| 3 | 00:20 | use | bookstack | disk 95% full |
|
||||
| 4 | 00:45 | offsite-run | mealie | tunnel down 10min |
|
||||
| 5 | 01:10 | use | privatebin | docker restarted |
|
||||
| 6 | 01:35 | backup-system | adventurelog | nothing |
|
||||
| 7 | 02:00 | update | nextcloud | internet gone 10min |
|
||||
| 8 | 02:25 | backup-app | nextcloud | internet gone 10min |
|
||||
| 9 | 02:50 | use | uptime-kuma | internet gone 10min |
|
||||
| 10 | 03:15 | restore | uptime-kuma | hard reset |
|
||||
| 11 | 03:40 | use | paperless-ngx | drive pulled 20min |
|
||||
| 12 | 04:05 | use | paperless-ngx | nothing |
|
||||
|
||||
**Re-draw log** — a silent re-draw is a schedule chosen by the person running it, so every one is here:
|
||||
|
||||
- r02 X=reinstall re-drawn (nothing has been removed yet)
|
||||
- r06 Z=disk 95% full re-drawn (constraint 4: at most once)
|
||||
- r07 Z=nothing re-drawn (constraint 6: never two in a row after r2)
|
||||
- r08 X=reinstall re-drawn (nothing has been removed yet)
|
||||
|
||||
**What the draw happened to give, said plainly before the night judges it:** no `remove` round was
|
||||
ever drawn, so `reinstall` had nothing to reinstall and was re-drawn twice (rounds 2 and 8). Three
|
||||
`internet gone` rounds land consecutively (7, 8, 9) — that is the seed's doing, and it makes rounds
|
||||
7–9 a de-facto endurance test of the same accident against three different actions rather than three
|
||||
independent samples. `controller killed`, `drive pulled 90s`, `memory pressure`, `agent restarted`
|
||||
and `hub unreachable` were never drawn at all; **this night does not test them**, and the morning
|
||||
verdict must not claim it did.
|
||||
|
||||
## Phase 0 — the golden, the box, the household
|
||||
|
||||
**0.1 Golden 0.245.0, baked and published.** Launched 19:52:47Z as a transient unit in the drill VM,
|
||||
finished 19:58:15Z. Markers: `overlay2`=1, `including mount point`=2, `upload OK (HTTP 201)`=1,
|
||||
FATAL=0, publish-skipped=0. `GOLDEN_SHA256=7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626`.
|
||||
**The teardown was gated on the REGISTRY answering 200**, not on an exit code — and that mattered:
|
||||
the wrapper exited **144** while every measured outcome was good. Vouched in the hub and read back
|
||||
from the page (`golden currently vouched: 0.245.0`). The three bake failures of 2026-09-16 (`scp -P`,
|
||||
`chmod 0700`, `GITEA_USER=admin`) were each guarded and none recurred.
|
||||
|
||||
*Decision, recorded because silence reads as agreement:* the global controller floor was left at
|
||||
0.244.0. The new box installs golden 0.245.0, which already carries controller 0.245.0, so no floor
|
||||
was needed to deliver anything tonight; raising it would have pushed an update onto demo-felhom, a
|
||||
box not in this drill.
|
||||
|
||||
**0.2 The box.** VM 336 on demo-hp: 8 GiB, 4 cores, 32 G system + 100 G data disk on the NVMe at its
|
||||
root, booted from the **published** ISO 1.28.0. Boot order set in its own `qm set` (combining it
|
||||
silently yields `order=net0;ide2`). Install completion was judged **from the disk** — blocks used
|
||||
grew 3233 → 6942 MiB then held across three checks — because „Automatically reboot" is ticked and a
|
||||
finished install looks exactly like a stuck one on screen. The summary page was read before pressing
|
||||
Install, and the line that made it safe was **„Disk(s): /dev/sda"** — the 32 G system disk alone.
|
||||
|
||||
**The walk, as a volunteer, cost ZERO operator presses.** The box registered itself as an unclaimed
|
||||
appliance and polled, visibly, until bound. The bind link came from the **waiting mail** (minted
|
||||
18:17:46Z by yesterday's acknowledged host delete), the pairing code off the box's own console
|
||||
(`4SY-4TX`), the „Tulajdonosi jelmondat" from the hub's customer record:
|
||||
POST /bind/<token> -> 200, „Sikeres összekötés."
|
||||
Then day-0 ran on its own and the hub recorded, without anyone pressing anything:
|
||||
`appliance_bound` (customer_selfbind) · `appliance_credential_delivered` · `claim_reissued_reenroll`
|
||||
· `offsite_reissued` · **`pbsdr_auto_reissue` — „Previous key destroyed (acknowledged deletion) —
|
||||
credentials re-issued automatically."**
|
||||
|
||||
**That last event is a first.** The brief named „the WG hook provisions by itself after an
|
||||
acknowledged delete" as a claim never measured live. It ran tonight, unprompted. **Both pre-declared
|
||||
presses (O1 self-bind, O2 re-issue) were therefore unnecessary.**
|
||||
|
||||
The dashboard was claimed with the mailed code and **proven by logging in with the new password** —
|
||||
a claim page that re-renders looks identical to success from the status code alone. The 100 GB data
|
||||
drive was initialised through the wizard's own endpoint (`POST /api/storage/init`, polled to
|
||||
`phase: done`), and `df` shows it mounted at `/mnt/felhom-drives/hdd_1` with 93 G free.
|
||||
|
||||
**The box landed on tonight's golden with no hand upgrade:** controller **0.245.0** (healthy), agent
|
||||
**0.131.0**, host `tester-1-022354` ONLINE. And the **R-543 escrow reminder bar shipped hours earlier
|
||||
was live on it**, on a box nobody had touched.
|
||||
|
||||
**0.3 The household — and the first real trouble.** Twelve deploys were fired; **ten were accepted,
|
||||
two refused** for memory with both numbers quoted. Then nine of the ten failed: the guest's disks are
|
||||
thin-provisioned over an ~11.8 GB pool carved from a 32 GB system disk, ten simultaneous image pulls
|
||||
filled it, and the hub recorded `storage_fill_critical` (100 %) plus **nine `app_deploy_failed`
|
||||
warnings**, one per app, each naming the failing pull. Only PrivateBin installed.
|
||||
|
||||
**The product behaved; the harness did not.** R-536's failure event — shipped that same morning so an
|
||||
interrupted install is not silence — fired for all nine within two minutes. The memory guard refused
|
||||
rather than over-committing. The per-stack record stayed honest (`deployed: false`). The two faults
|
||||
were mine: firing twelve deploys in two seconds is not household behaviour, and a 32 GB system disk
|
||||
was copied from an earlier drill without checking what that drill had installed.
|
||||
|
||||
**0.4 The escrow ceremony could NOT be completed — and this one is about tonight's own release.**
|
||||
See „Finding: the recovery-code step cannot be done when the guide says to do it" below.
|
||||
|
||||
|
||||
## Finding: the recovery-code step cannot be done when the guide says to do it (R-546)
|
||||
|
||||
The box was at exactly the point of the guide this release added hours earlier — installed, bound
|
||||
with no press, claimed, drive initialised, **no apps yet** — and the escrow reminder bar was on every
|
||||
page telling the household to create their recovery code. It could not be done.
|
||||
|
||||
POST /api/escrow/start -> 200 {"job_id":"escrow-1789590499667361664","phase":"running"}
|
||||
GET /api/escrow/status -> claimable:false —
|
||||
detail: "exit 2: … selftest=escrow-create requires -storage <pbs-storage-id> (or escrow.pbs_storage…"
|
||||
POST /api/escrow/claim -> **409** „A folyamat jelenlegi állapotában a kód nem kérhető le."
|
||||
|
||||
**Both sides agreed on the cause.** The hub's own Backup & DR panel read „host enrolled **done** · WG
|
||||
tunnel peer registered **done** · descriptor provisioned (namespace tester-1, token felhom@pbs!tester-1)
|
||||
**waiting** · ceremony possible once the descriptor is applied on the box". The box had no PBS storage
|
||||
(`pvesm status`: `local`, `local-lvm` only) and no `escrow` section in `agent.json` at all.
|
||||
|
||||
**It self-heals, and that was measured rather than assumed.** The box was left alone and polled:
|
||||
|
||||
20:30:12Z … 20:34:15Z pbs_storage=none escrow.pbs_storage_id=none
|
||||
**20:35:16Z pbs_storage=felhom-pbs escrow.pbs_storage_id=felhom-pbs**
|
||||
|
||||
~17 minutes after the bind. The retried ceremony passed every preflight item, the claim returned
|
||||
**200** (83-character code, 129.2 bits of entropy, revealed once), and `escrow_state` became
|
||||
**escrowed**. The bar then vanished from all four pages checked — the R-543 fix working through its
|
||||
whole lifecycle on a box nobody had set up for the test.
|
||||
|
||||
**So the defect is timing and wording, not mechanism.** For ~17 minutes a volunteer following
|
||||
tonight's guide meets a stderr fragment about a `-storage` flag, while every page urges them on.
|
||||
Filed **R-546** (P2). No product code was changed — this is a validation run.
|
||||
|
||||
## Phase 1 — the twelve rounds
|
||||
|
||||
PENDING
|
||||
|
||||
## Phase 2 — the morning after
|
||||
|
||||
PENDING
|
||||
|
||||
## Interventions — counted, with the reason for each verdict
|
||||
|
||||
PENDING
|
||||
|
||||
## Teardown — three layers, stated
|
||||
|
||||
PENDING
|
||||
@@ -0,0 +1,325 @@
|
||||
[golden] build-golden.sh v3.0.0 — baking controller gitea.dooplex.hu/admin/felhom-controller:0.245.0
|
||||
[golden] creating build LXC 9100 (nesting=1,keyctl=1, unprivileged; rootfs 32G + ONE data volume 24G @ /var/lib/felhom, backup=1) …
|
||||
Logical volume "vm-9100-disk-0" created.
|
||||
Logical volume pve/vm-9100-disk-0 changed.
|
||||
Creating filesystem with 8388608 4k blocks and 2097152 inodes
|
||||
Filesystem UUID: 7b98df7b-8d9d-43c8-ba99-0ba1075e83cc
|
||||
Superblock backups stored on blocks:
|
||||
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
|
||||
4096000, 7962624
|
||||
Logical volume "vm-9100-disk-1" created.
|
||||
Logical volume pve/vm-9100-disk-1 changed.
|
||||
Creating filesystem with 6291456 4k blocks and 1572864 inodes
|
||||
Filesystem UUID: e121482f-10e3-4687-b082-84db662425f4
|
||||
Superblock backups stored on blocks:
|
||||
32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
|
||||
extracting archive '/var/lib/vz/template/cache/debian-13-standard_13.6-1_amd64.tar.zst'
|
||||
Total bytes read: 553512960 (528MiB, 86MiB/s)
|
||||
Detected container architecture: amd64
|
||||
Creating SSH host key 'ssh_host_ecdsa_key' - this may take some time ...
|
||||
done: SHA256:aX2YgyXohsbFO+/bVcsxaM32nSEJWtuQRE7krfEzqQ0 root@felhom-golden
|
||||
Creating SSH host key 'ssh_host_ed25519_key' - this may take some time ...
|
||||
done: SHA256:QwivNTDkZ0thSB8KcWpCc0GTZ5He3NIFgGDmWBD5DQQ root@felhom-golden
|
||||
Creating SSH host key 'ssh_host_rsa_key' - this may take some time ...
|
||||
done: SHA256:HDmBrODCYm0syAs+LuleMy+QMBc/CA/GIPdmH0bppJA root@felhom-golden
|
||||
[golden] starting + installing Docker (official repo, trixie channel) …
|
||||
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = (unset),
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to the standard locale ("C").
|
||||
locale: Cannot set LC_CTYPE to default locale: No such file or directory
|
||||
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
|
||||
locale: Cannot set LC_ALL to default locale: No such file or directory
|
||||
apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct!
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = (unset),
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to the standard locale ("C").
|
||||
locale: Cannot set LC_CTYPE to default locale: No such file or directory
|
||||
locale: Cannot set LC_MESSAGES to default locale: No such file or directory
|
||||
locale: Cannot set LC_ALL to default locale: No such file or directory
|
||||
[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation …
|
||||
[golden] wiring the single data volume (R-165 variant V-c): /var/lib/felhom/{docker,sys_drive} -> binds …
|
||||
[golden] verifying Docker works in the build guest (storage driver should be overlay2 on the ext4 data volume) …
|
||||
Unable to find image 'hello-world:latest' locally
|
||||
latest: Pulling from library/hello-world
|
||||
4f55086f7dd0: Pulling fs layer
|
||||
4f55086f7dd0: Verifying Checksum
|
||||
4f55086f7dd0: Download complete
|
||||
4f55086f7dd0: Pull complete
|
||||
Digest: sha256:5e23090353324d887c48ad5e5c56d294eab81588df9605b07d1afe895f9cc8f8
|
||||
Status: Downloaded newer image for hello-world:latest
|
||||
docker OK (overlay2; data-root /var/lib/docker)
|
||||
/var/lib/docker is a real mount: /dev/mapper/pve-vm--9100--disk--1[/docker] ext4
|
||||
/mnt/sys_drive is a real mount: /dev/mapper/pve-vm--9100--disk--1[/sys_drive] ext4
|
||||
both paths are ONE filesystem: /dev/mapper/pve-vm--9100--disk--1 23317576
|
||||
[golden] baking the in-guest controller image gitea.dooplex.hu/admin/felhom-controller:0.245.0 (no registry cred at deploy) …
|
||||
|
||||
WARNING! Your credentials are stored unencrypted in '/root/.docker/config.json'.
|
||||
Configure a credential helper to remove this warning. See
|
||||
https://docs.docker.com/go/credential-store/
|
||||
|
||||
0.245.0: Pulling from admin/felhom-controller
|
||||
a8ac7f6c67ab: Pulling fs layer
|
||||
bf30769d36e7: Pulling fs layer
|
||||
044b66fbe46c: Pulling fs layer
|
||||
b5c41a28e83f: Pulling fs layer
|
||||
965b73d03024: Pulling fs layer
|
||||
9dd06928817a: Pulling fs layer
|
||||
b5c41a28e83f: Waiting
|
||||
965b73d03024: Waiting
|
||||
9dd06928817a: Waiting
|
||||
044b66fbe46c: Verifying Checksum
|
||||
044b66fbe46c: Download complete
|
||||
b5c41a28e83f: Verifying Checksum
|
||||
b5c41a28e83f: Download complete
|
||||
a8ac7f6c67ab: Verifying Checksum
|
||||
a8ac7f6c67ab: Download complete
|
||||
965b73d03024: Verifying Checksum
|
||||
965b73d03024: Download complete
|
||||
9dd06928817a: Verifying Checksum
|
||||
9dd06928817a: Download complete
|
||||
bf30769d36e7: Verifying Checksum
|
||||
bf30769d36e7: Download complete
|
||||
a8ac7f6c67ab: Pull complete
|
||||
bf30769d36e7: Pull complete
|
||||
044b66fbe46c: Pull complete
|
||||
b5c41a28e83f: Pull complete
|
||||
965b73d03024: Pull complete
|
||||
9dd06928817a: Pull complete
|
||||
Digest: sha256:b6abd24d67be8ef1aa61b30f852f3a1f6e90ec3d3599a3824655a290c0dba875
|
||||
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-controller:0.245.0
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.245.0
|
||||
[golden] asking the controller which infra images it manages …
|
||||
[golden] baking infra images (4): traefik:v3.6.7 cloudflare/cloudflared:2026.6.0 gtstef/filebrowser:1.3.3-stable gitea.dooplex.hu/admin/felhom-samba:1.1.0 …
|
||||
v3.6.7: Pulling from library/traefik
|
||||
589002ba0eae: Pulling fs layer
|
||||
ef63511ea6cc: Pulling fs layer
|
||||
0738e5cb835e: Pulling fs layer
|
||||
3e6813f70c64: Pulling fs layer
|
||||
3e6813f70c64: Waiting
|
||||
589002ba0eae: Verifying Checksum
|
||||
589002ba0eae: Download complete
|
||||
ef63511ea6cc: Verifying Checksum
|
||||
ef63511ea6cc: Download complete
|
||||
3e6813f70c64: Verifying Checksum
|
||||
3e6813f70c64: Download complete
|
||||
589002ba0eae: Pull complete
|
||||
0738e5cb835e: Verifying Checksum
|
||||
0738e5cb835e: Download complete
|
||||
ef63511ea6cc: Pull complete
|
||||
0738e5cb835e: Pull complete
|
||||
3e6813f70c64: Pull complete
|
||||
Digest: sha256:a9890c898f379c1905ee5b28342f6b408dc863f08db2dab20e46c267d1ff463a
|
||||
Status: Downloaded newer image for traefik:v3.6.7
|
||||
docker.io/library/traefik:v3.6.7
|
||||
2026.6.0: Pulling from cloudflare/cloudflared
|
||||
47de5dd0b812: Pulling fs layer
|
||||
c172f21841df: Pulling fs layer
|
||||
99515e7b4d35: Pulling fs layer
|
||||
99ba982a9142: Pulling fs layer
|
||||
d6b1b89eccac: Pulling fs layer
|
||||
2780920e5dbf: Pulling fs layer
|
||||
7c12895b777b: Pulling fs layer
|
||||
3214acf345c0: Pulling fs layer
|
||||
52630fc75a18: Pulling fs layer
|
||||
dd64bf2dd177: Pulling fs layer
|
||||
b839dfae01f6: Pulling fs layer
|
||||
ebddc55facdc: Pulling fs layer
|
||||
bdfd7f7e5bf6: Pulling fs layer
|
||||
2d4d7adf6272: Pulling fs layer
|
||||
40008157d8d2: Pulling fs layer
|
||||
bd8962e29291: Pulling fs layer
|
||||
cac2ae0193cb: Pulling fs layer
|
||||
74d1dac84ecc: Pulling fs layer
|
||||
dd64bf2dd177: Waiting
|
||||
b839dfae01f6: Waiting
|
||||
ebddc55facdc: Waiting
|
||||
bdfd7f7e5bf6: Waiting
|
||||
2d4d7adf6272: Waiting
|
||||
40008157d8d2: Waiting
|
||||
bd8962e29291: Waiting
|
||||
cac2ae0193cb: Waiting
|
||||
74d1dac84ecc: Waiting
|
||||
2780920e5dbf: Waiting
|
||||
7c12895b777b: Waiting
|
||||
3214acf345c0: Waiting
|
||||
52630fc75a18: Waiting
|
||||
99ba982a9142: Waiting
|
||||
d6b1b89eccac: Waiting
|
||||
c172f21841df: Download complete
|
||||
47de5dd0b812: Verifying Checksum
|
||||
99515e7b4d35: Verifying Checksum
|
||||
99515e7b4d35: Download complete
|
||||
99ba982a9142: Verifying Checksum
|
||||
99ba982a9142: Download complete
|
||||
d6b1b89eccac: Verifying Checksum
|
||||
d6b1b89eccac: Download complete
|
||||
47de5dd0b812: Pull complete
|
||||
2780920e5dbf: Verifying Checksum
|
||||
2780920e5dbf: Download complete
|
||||
7c12895b777b: Verifying Checksum
|
||||
7c12895b777b: Download complete
|
||||
3214acf345c0: Verifying Checksum
|
||||
3214acf345c0: Download complete
|
||||
52630fc75a18: Verifying Checksum
|
||||
52630fc75a18: Download complete
|
||||
dd64bf2dd177: Verifying Checksum
|
||||
dd64bf2dd177: Download complete
|
||||
b839dfae01f6: Verifying Checksum
|
||||
b839dfae01f6: Download complete
|
||||
c172f21841df: Pull complete
|
||||
ebddc55facdc: Verifying Checksum
|
||||
ebddc55facdc: Download complete
|
||||
bdfd7f7e5bf6: Verifying Checksum
|
||||
bdfd7f7e5bf6: Download complete
|
||||
2d4d7adf6272: Verifying Checksum
|
||||
2d4d7adf6272: Download complete
|
||||
bd8962e29291: Verifying Checksum
|
||||
bd8962e29291: Download complete
|
||||
cac2ae0193cb: Verifying Checksum
|
||||
cac2ae0193cb: Download complete
|
||||
74d1dac84ecc: Verifying Checksum
|
||||
74d1dac84ecc: Download complete
|
||||
40008157d8d2: Verifying Checksum
|
||||
40008157d8d2: Download complete
|
||||
99515e7b4d35: Pull complete
|
||||
99ba982a9142: Pull complete
|
||||
d6b1b89eccac: Pull complete
|
||||
2780920e5dbf: Pull complete
|
||||
7c12895b777b: Pull complete
|
||||
3214acf345c0: Pull complete
|
||||
52630fc75a18: Pull complete
|
||||
dd64bf2dd177: Pull complete
|
||||
b839dfae01f6: Pull complete
|
||||
ebddc55facdc: Pull complete
|
||||
bdfd7f7e5bf6: Pull complete
|
||||
2d4d7adf6272: Pull complete
|
||||
40008157d8d2: Pull complete
|
||||
bd8962e29291: Pull complete
|
||||
cac2ae0193cb: Pull complete
|
||||
74d1dac84ecc: Pull complete
|
||||
Digest: sha256:ba461b8aa9c042156dbd39c38657fe7431bafa063220eab8d5330a523863da9f
|
||||
Status: Downloaded newer image for cloudflare/cloudflared:2026.6.0
|
||||
docker.io/cloudflare/cloudflared:2026.6.0
|
||||
1.3.3-stable: Pulling from gtstef/filebrowser
|
||||
6a0ac1617861: Pulling fs layer
|
||||
ef8806083e82: Pulling fs layer
|
||||
b74107c861c7: Pulling fs layer
|
||||
adc935def003: Pulling fs layer
|
||||
4f4fb700ef54: Pulling fs layer
|
||||
18695ccc900a: Pulling fs layer
|
||||
45d119d5c397: Pulling fs layer
|
||||
dac52db4fc51: Pulling fs layer
|
||||
6d598f86b2f2: Pulling fs layer
|
||||
8aa349c8396c: Pulling fs layer
|
||||
adc935def003: Waiting
|
||||
4f4fb700ef54: Waiting
|
||||
18695ccc900a: Waiting
|
||||
45d119d5c397: Waiting
|
||||
dac52db4fc51: Waiting
|
||||
6d598f86b2f2: Waiting
|
||||
8aa349c8396c: Waiting
|
||||
6a0ac1617861: Verifying Checksum
|
||||
6a0ac1617861: Download complete
|
||||
b74107c861c7: Verifying Checksum
|
||||
b74107c861c7: Download complete
|
||||
adc935def003: Verifying Checksum
|
||||
adc935def003: Download complete
|
||||
4f4fb700ef54: Verifying Checksum
|
||||
4f4fb700ef54: Download complete
|
||||
45d119d5c397: Verifying Checksum
|
||||
45d119d5c397: Download complete
|
||||
ef8806083e82: Verifying Checksum
|
||||
ef8806083e82: Download complete
|
||||
dac52db4fc51: Verifying Checksum
|
||||
dac52db4fc51: Download complete
|
||||
6d598f86b2f2: Download complete
|
||||
18695ccc900a: Verifying Checksum
|
||||
18695ccc900a: Download complete
|
||||
6a0ac1617861: Pull complete
|
||||
8aa349c8396c: Download complete
|
||||
ef8806083e82: Pull complete
|
||||
b74107c861c7: Pull complete
|
||||
adc935def003: Pull complete
|
||||
4f4fb700ef54: Pull complete
|
||||
18695ccc900a: Pull complete
|
||||
45d119d5c397: Pull complete
|
||||
dac52db4fc51: Pull complete
|
||||
6d598f86b2f2: Pull complete
|
||||
8aa349c8396c: Pull complete
|
||||
Digest: sha256:eb3733681db8757412632c61a99ad656f0d94ed6781bb2ea114b4d70babab78c
|
||||
Status: Downloaded newer image for gtstef/filebrowser:1.3.3-stable
|
||||
docker.io/gtstef/filebrowser:1.3.3-stable
|
||||
1.1.0: Pulling from admin/felhom-samba
|
||||
897d797d2723: Pulling fs layer
|
||||
3051591aa250: Pulling fs layer
|
||||
ce57a3f93416: Pulling fs layer
|
||||
fb94eeec2fe1: Pulling fs layer
|
||||
fb94eeec2fe1: Waiting
|
||||
ce57a3f93416: Verifying Checksum
|
||||
ce57a3f93416: Download complete
|
||||
fb94eeec2fe1: Verifying Checksum
|
||||
fb94eeec2fe1: Download complete
|
||||
897d797d2723: Verifying Checksum
|
||||
897d797d2723: Download complete
|
||||
3051591aa250: Verifying Checksum
|
||||
3051591aa250: Download complete
|
||||
897d797d2723: Pull complete
|
||||
3051591aa250: Pull complete
|
||||
ce57a3f93416: Pull complete
|
||||
fb94eeec2fe1: Pull complete
|
||||
Digest: sha256:1c17c09422bec0366d7cf0e0fcfc1486ba6c90334a0a5d5c851073a9342f8f10
|
||||
Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-samba:1.1.0
|
||||
gitea.dooplex.hu/admin/felhom-samba:1.1.0
|
||||
[golden] baking the controller-bootstrap unit (deploys the BAKED controller from the config mount) …
|
||||
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.service' → '/etc/systemd/system/felhom-controller-bootstrap.service'.
|
||||
[golden] baking the controller-bootstrap PATH unit (starts the service on bootstrap-mount hot-plug — B1) …
|
||||
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.path' → '/etc/systemd/system/felhom-controller-bootstrap.path'.
|
||||
[golden] baking the first-boot SSH host-key regeneration unit (F3) …
|
||||
Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-regen-hostkeys.service' → '/etc/systemd/system/felhom-regen-hostkeys.service'.
|
||||
[golden] identity-clean + minimize …
|
||||
[golden] stop + archive …
|
||||
INFO: including mount point rootfs ('/') in backup
|
||||
INFO: including mount point mp0 ('/var/lib/felhom') in backup
|
||||
INFO: archive file size: 623MB
|
||||
INFO: Finished Backup of VM 9100 (00:00:31)
|
||||
[golden] DONE. golden archive volid: local:backup/vzdump-lxc-9100-2026_09_16-21_57_05.tar.zst (rootfs 32G + ONE data volume 24G @ /var/lib/felhom, all in the archive)
|
||||
[golden] publishing golden (653729820 bytes, sha256 7a08aa1ad0bdd622…) → https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.245.0/golden.tar.zst
|
||||
[golden] pre-delete existing: HTTP 404 (404/204 expected)
|
||||
[golden] upload OK (HTTP 201)
|
||||
GOLDEN_VERSION=0.245.0
|
||||
GOLDEN_SHA256=7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626
|
||||
[golden] Record in the hub operator UI (Configs → Day-0 artifacts): golden 0.245.0 / 7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626
|
||||
[golden] (the build guest 9100 is stopped; destroy it with: pct destroy 9100 --purge)
|
||||
@@ -0,0 +1,154 @@
|
||||
#!/usr/bin/env python3
|
||||
# -*- coding: utf-8 -*-
|
||||
"""chaos_schedule.py — draw the CHAOS NIGHT schedule, deterministically, ONCE.
|
||||
|
||||
The point of this file is that the night cannot be chosen after the fact. The seed is the date, the
|
||||
draws are stdlib `random` seeded with it, and the table this prints goes into the findings document
|
||||
BEFORE round 1 runs. Anyone can re-run it and get the same night.
|
||||
|
||||
A draw that breaks a constraint is RE-DRAWN and the re-draw is logged, because a silent re-draw is a
|
||||
schedule chosen by the person running it.
|
||||
"""
|
||||
import hashlib
|
||||
import random
|
||||
import sys
|
||||
|
||||
SEED = 20260917
|
||||
|
||||
ACTIONS = [ # (name, weight, what a household does)
|
||||
("use", 3, "ten minutes in the app: create, edit, upload, delete one thing"),
|
||||
("backup-app", 2, "„Mentés most” on the app"),
|
||||
("backup-system", 1, "whole-system „Mentés most”"),
|
||||
("offsite-run", 1, "the tier-3 leg through its own endpoint"),
|
||||
("update", 1, "the guarded Update on the app the catalog bumped"),
|
||||
("restore", 1, "restore the app through the page (wizard for off-site)"),
|
||||
("remove", 1, "remove the app with „az adataimat is töröld”"),
|
||||
("reinstall", 1, "reinstall an app removed earlier tonight, and restore it"),
|
||||
]
|
||||
|
||||
APPS = ["nextcloud", "immich", "bookstack", "privatebin", "gokapi", "vaultwarden",
|
||||
"paperless-ngx", "jellyfin", "mealie", "uptime-kuma", "adventurelog", "homebox"]
|
||||
# remove/reinstall may only touch these, so the photo and document apps survive for the morning restore
|
||||
DISPOSABLE = ["privatebin", "gokapi", "homebox", "mealie"]
|
||||
|
||||
ACCIDENTS = [ # (name, weight, how)
|
||||
("nothing", 3, "—"),
|
||||
("power cut", 2, "qm stop, 60 s, qm start"),
|
||||
("hard reset", 1, "qm reset"),
|
||||
("drive pulled 90s", 1, "detach the data disk, reattach after 90 s"),
|
||||
("drive pulled 20min", 1, "detach the data disk, reattach after 20 min"),
|
||||
("internet gone 10min", 2, "block the VM's outbound at the host, LAN kept"),
|
||||
("hub unreachable 15min", 1, "block only the hub's address from the VM"),
|
||||
("controller killed", 1, "docker kill felhom-controller"),
|
||||
("agent restarted", 1, "systemctl restart felhom-agent in the nested PVE"),
|
||||
("docker restarted", 1, "systemctl restart docker in the guest"),
|
||||
("disk 95% full", 1, "fill the system disk to 95 % for 10 min, then free it"),
|
||||
("tunnel down 10min", 1, "docker kill cloudflared"),
|
||||
("memory pressure", 1, "a throwaway container with a 1 GB hog for 5 min"),
|
||||
]
|
||||
|
||||
ROUNDS = 12
|
||||
START_MIN = 23 * 60 + 30 # 23:30
|
||||
SPACING = 25
|
||||
|
||||
|
||||
def wpick(rng, table):
|
||||
names = [t[0] for t in table]
|
||||
weights = [t[1] for t in table]
|
||||
return rng.choices(names, weights=weights, k=1)[0]
|
||||
|
||||
|
||||
def hhmm(total):
|
||||
total %= 24 * 60
|
||||
return "%02d:%02d" % (total // 60, total % 60)
|
||||
|
||||
|
||||
def draw():
|
||||
rng = random.Random(SEED)
|
||||
rows, log = [], []
|
||||
used = {"update": 0, "controller killed": 0, "drive pulled 20min": 0, "disk 95% full": 0}
|
||||
removed = [] # apps removed earlier tonight (reinstall needs one)
|
||||
prev_accident = None
|
||||
|
||||
for n in range(1, ROUNDS + 1):
|
||||
t = hhmm(START_MIN + (n - 1) * SPACING)
|
||||
|
||||
# ---- X, the action -------------------------------------------------
|
||||
for attempt in range(1, 40):
|
||||
x = wpick(rng, ACTIONS)
|
||||
if x == "update" and used["update"] >= 1:
|
||||
log.append("r%02d X=update -> becomes 'use' (constraint 3: there is one bump)" % n)
|
||||
x = "use"
|
||||
if x == "reinstall" and not removed:
|
||||
log.append("r%02d X=reinstall re-drawn (nothing has been removed yet)" % n)
|
||||
continue
|
||||
break
|
||||
|
||||
# ---- Y, the app ----------------------------------------------------
|
||||
if x == "remove":
|
||||
pool = [a for a in DISPOSABLE if a not in removed]
|
||||
if not pool:
|
||||
log.append("r%02d X=remove -> becomes 'use' (every disposable app is already removed)" % n)
|
||||
x, pool = "use", APPS
|
||||
y = rng.choice(pool)
|
||||
elif x == "reinstall":
|
||||
y = rng.choice(removed)
|
||||
else:
|
||||
y = rng.choice(APPS)
|
||||
|
||||
# ---- Z, the accident ----------------------------------------------
|
||||
if n in (1, ROUNDS):
|
||||
z = "nothing" # constraint 5: control rounds
|
||||
else:
|
||||
for attempt in range(1, 60):
|
||||
z = wpick(rng, ACCIDENTS)
|
||||
if z == "nothing" and prev_accident == "nothing" and n > 2:
|
||||
log.append("r%02d Z=nothing re-drawn (constraint 6: never two in a row after r2)" % n)
|
||||
continue
|
||||
if z == "controller killed":
|
||||
if used["controller killed"] >= 3:
|
||||
log.append("r%02d Z=controller killed re-drawn (constraint 2: max 3 a night)" % n)
|
||||
continue
|
||||
if x == "restore":
|
||||
log.append("r%02d Z=controller killed re-drawn (constraint 2: never with 'restore' — the F9/F3 class is already measured)" % n)
|
||||
continue
|
||||
if z == "drive pulled 20min" and used["drive pulled 20min"] >= 1:
|
||||
log.append("r%02d Z=drive pulled 20min re-drawn (constraint 4: at most once)" % n)
|
||||
continue
|
||||
if z == "disk 95% full" and used["disk 95% full"] >= 1:
|
||||
log.append("r%02d Z=disk 95%% full re-drawn (constraint 4: at most once)" % n)
|
||||
continue
|
||||
break
|
||||
|
||||
if x in used:
|
||||
used[x] += 1
|
||||
if z in used:
|
||||
used[z] += 1
|
||||
if x == "remove":
|
||||
removed.append(y)
|
||||
if x == "reinstall" and y in removed:
|
||||
removed.remove(y)
|
||||
prev_accident = z
|
||||
rows.append((n, t, x, y, z))
|
||||
|
||||
return rows, log
|
||||
|
||||
|
||||
def main():
|
||||
rows, log = draw()
|
||||
h = hashlib.sha256(open(__file__, "rb").read()).hexdigest()
|
||||
print("seed: %d" % SEED)
|
||||
print("script sha256: %s" % h)
|
||||
print()
|
||||
print("| # | time | X — the action | Y — the app | Z — the accident |")
|
||||
print("|---|---|---|---|---|")
|
||||
for n, t, x, y, z in rows:
|
||||
print("| %d | %s | %s | %s | %s |" % (n, t, x, y, z))
|
||||
print()
|
||||
print("re-draw log (%d entries):" % len(log))
|
||||
for line in log or [" (none)"]:
|
||||
print(" " + line)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -0,0 +1,33 @@
|
||||
#!/bin/bash
|
||||
# events.sh — dump the hub's recorded events for the drill customer.
|
||||
#
|
||||
# This is the alarm-scoring surface. It reads the customer page's `events-table`, which is the
|
||||
# hub's own record — NOT my notes. The hub pod has no sqlite3 and the DB is 300 MB, so the page is
|
||||
# the right instrument; a copied hub.db without its -wal reads hours stale here.
|
||||
#
|
||||
# Usage: events.sh [n] — print the newest n rows (default 25)
|
||||
set -u
|
||||
N="${1:-25}"
|
||||
cd /mnt/5_hdd/felhom.eu/git/felhom.eu
|
||||
python3 scripts/read_credential.py HUB_PW /tmp/.ehp >/dev/null 2>&1
|
||||
HUB_PW=$(cat /tmp/.ehp); IP=$(sudo kubectl -n felhom-system get svc hub -o jsonpath='{.spec.clusterIP}')
|
||||
curl -s -u ":$HUB_PW" "http://$IP:8080/customers/tester-1" -o /tmp/.ev.html
|
||||
rm -f /tmp/.ehp
|
||||
python3 - "$N" <<'PY'
|
||||
import re,html,sys
|
||||
n=int(sys.argv[1])
|
||||
t=open('/tmp/.ev.html',encoding='utf-8',errors='replace').read()
|
||||
m=re.search(r'id="events-table"(.*?)</table>', t, flags=re.S)
|
||||
if not m:
|
||||
print("EVENTS TABLE NOT FOUND — the instrument failed, and that is not the same as 'no events'")
|
||||
raise SystemExit(1)
|
||||
rows=re.findall(r'<tr[^>]*>(.*?)</tr>', m.group(1), flags=re.S)
|
||||
out=0
|
||||
for r in rows:
|
||||
cells=[html.unescape(re.sub(r'\s+',' ',re.sub(r'<[^>]+>','',c))).strip() for c in re.findall(r'<t[dh][^>]*>(.*?)</t[dh]>', r, flags=re.S)]
|
||||
if not cells: continue
|
||||
print(" | " + " | ".join(cells)[:200])
|
||||
out+=1
|
||||
if out>n: break
|
||||
PY
|
||||
rm -f /tmp/.ev.html
|
||||
@@ -0,0 +1,36 @@
|
||||
#!/bin/bash
|
||||
# household_loop.sh — the light background household (CHAOS NIGHT Phase 0.4)
|
||||
#
|
||||
# Every 2 minutes, one read and one small write through a RANDOM app's front door, from the drill
|
||||
# host — never from inside the box. It logs `time app op result`. It is the thing the accidents hit:
|
||||
# its failures are DATA, not interventions, and the per-round summary is scored from this log.
|
||||
#
|
||||
# Usage: household_loop.sh <box-ip> <logfile>
|
||||
set -u
|
||||
BOX="${1:?box ip}"
|
||||
LOG="${2:?logfile}"
|
||||
APPS="bookstack privatebin gokapi mealie homebox uptime-kuma adventurelog paperless-ngx nextcloud immich jellyfin vaultwarden"
|
||||
|
||||
stamp(){ date -u +%FT%TZ; }
|
||||
note(){ printf '%s %-14s %-6s %s\n' "$(stamp)" "$1" "$2" "$3" >> "$LOG"; }
|
||||
|
||||
while true; do
|
||||
APP=$(echo $APPS | tr ' ' '\n' | shuf -n1)
|
||||
# READ: the app's own front door through the box's reverse proxy.
|
||||
CODE=$(curl -s -o /dev/null -w '%{http_code}' --max-time 15 \
|
||||
-H "Host: ${APP}.enkicsifelhom.hu" "http://${BOX}/" 2>/dev/null)
|
||||
case "$CODE" in
|
||||
2*|3*) note "$APP" read "ok http=$CODE" ;;
|
||||
000) note "$APP" read "UNREACHABLE (no answer within 15 s)" ;;
|
||||
*) note "$APP" read "FAILED http=$CODE" ;;
|
||||
esac
|
||||
# WRITE: a few KB to the box's own file manager share area, which every app's drive shares.
|
||||
W=$(curl -s -o /dev/null -w '%{http_code}' --max-time 15 \
|
||||
-H "Host: felhom.enkicsifelhom.hu" "http://${BOX}/api/health" 2>/dev/null)
|
||||
case "$W" in
|
||||
2*) note "$APP" write "ok http=$W" ;;
|
||||
000) note "$APP" write "UNREACHABLE" ;;
|
||||
*) note "$APP" write "FAILED http=$W" ;;
|
||||
esac
|
||||
sleep 120
|
||||
done
|
||||
@@ -0,0 +1,76 @@
|
||||
#!/bin/bash
|
||||
# inject.sh — CHAOS NIGHT accident injectors. Run from DooPlex.
|
||||
#
|
||||
# Only the seven accidents the seed actually drew are implemented. The other six in the brief's
|
||||
# table were never drawn, and writing injectors for them would suggest this night tested them.
|
||||
#
|
||||
# Usage: inject.sh <accident> <round>
|
||||
# power-cut | hard-reset | drive-pulled-20min | internet-gone-10min
|
||||
# tunnel-down-10min | docker-restart | disk-95-full
|
||||
set -u
|
||||
A="${1:?accident}"; R="${2:?round}"
|
||||
VM=336
|
||||
E=/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17
|
||||
LOG="$E/round-${R}-accident.txt"
|
||||
H(){ ssh hp "$@" 2>/dev/null | grep -vE "locale|LC_|LANG|perl:|supported and installed|are supported"; }
|
||||
say(){ echo "$(date -u +%FT%TZ) $*" | tee -a "$LOG"; }
|
||||
# the customer guest lives INSIDE the nested PVE; these run one hop further in
|
||||
G(){ H "qm guest exec $VM -- $*" ; }
|
||||
|
||||
say "ACCIDENT=$A round=$R"
|
||||
case "$A" in
|
||||
power-cut)
|
||||
say "qm stop $VM (the plug is pulled — no clean shutdown)"
|
||||
H "qm stop $VM"; say "stopped; waiting 60 s with the box dark"
|
||||
sleep 60
|
||||
H "qm start $VM"; say "power back on"
|
||||
;;
|
||||
hard-reset)
|
||||
say "qm reset $VM (the reset button, mid-write)"
|
||||
H "qm reset $VM"; say "reset issued"
|
||||
;;
|
||||
drive-pulled-20min)
|
||||
F=$(H "qm config $VM | grep '^scsi1:' | sed 's/scsi1: //; s/,.*//'")
|
||||
say "data disk is $F — detaching it from the RUNNING box (the cable is pulled)"
|
||||
H "qm set $VM --delete scsi1"; say "detached; the disk file stays as unused0"
|
||||
say "leaving it out for 20 minutes"
|
||||
sleep 1200
|
||||
H "qm set $VM --scsi1 $F"; say "re-attached: $F"
|
||||
;;
|
||||
internet-gone-10min)
|
||||
TAP=$(H "ls /sys/class/net | grep -E \"^tap${VM}i0$\"")
|
||||
[ -n "$TAP" ] || { say "NO TAP FOUND for VM $VM — accident NOT injected, and that is recorded as such"; exit 1; }
|
||||
say "blocking the box's traffic off-LAN at the HOST, on $TAP; the LAN stays up"
|
||||
H "sysctl -w net.bridge.bridge-nf-call-iptables=1 >/dev/null
|
||||
iptables -I FORWARD 1 -m physdev --physdev-in $TAP -d 192.168.0.0/24 -j ACCEPT
|
||||
iptables -I FORWARD 2 -m physdev --physdev-in $TAP -j DROP"
|
||||
say "blocked (LAN allowed, everything else dropped) — 10 minutes"
|
||||
sleep 600
|
||||
H "iptables -D FORWARD -m physdev --physdev-in $TAP -j DROP
|
||||
iptables -D FORWARD -m physdev --physdev-in $TAP -d 192.168.0.0/24 -j ACCEPT
|
||||
sysctl -w net.bridge.bridge-nf-call-iptables=0 >/dev/null"
|
||||
say "unblocked; host sysctl restored to 0 and both rules removed"
|
||||
H "iptables -S FORWARD | head -5" | tee -a "$LOG"
|
||||
;;
|
||||
tunnel-down-10min)
|
||||
say "docker kill cloudflared inside the customer guest"
|
||||
G "docker kill cloudflared" ; say "tunnel killed — 10 minutes"
|
||||
sleep 600
|
||||
say "10 minutes up; NOT restarting it by hand — whether it returns by itself IS the measurement"
|
||||
;;
|
||||
docker-restart)
|
||||
say "systemctl restart docker inside the customer guest"
|
||||
G "systemctl restart docker"; say "docker restarted"
|
||||
;;
|
||||
disk-95-full)
|
||||
say "filling the customer guest's SYSTEM disk to ~95 %"
|
||||
G "bash -c 'df -h / | tail -1'" | tee -a "$LOG"
|
||||
G "bash -c 'F=\$(df --output=avail -m / | tail -1); fallocate -l \$(( (F * 95 / 100) ))M /var/tmp/.chaosfill && df -h / | tail -1'" | tee -a "$LOG"
|
||||
say "full — holding 10 minutes"
|
||||
sleep 600
|
||||
G "bash -c 'rm -f /var/tmp/.chaosfill; df -h / | tail -1'" | tee -a "$LOG"
|
||||
say "freed"
|
||||
;;
|
||||
*) say "UNKNOWN ACCIDENT $A"; exit 2 ;;
|
||||
esac
|
||||
say "accident $A complete"
|
||||
@@ -0,0 +1,34 @@
|
||||
#!/bin/bash
|
||||
# morning_after.sh — CHAOS NIGHT Phase 2. Run at ~05:00.
|
||||
#
|
||||
# Four things, in this order, and each one asks the product rather than my notes:
|
||||
# 1. every app healthy through its FRONT DOOR; every version label true; every backup page honest
|
||||
# 2. one DB-backed app restored from the OFF-SITE tier onto scratch guest 9202, read back
|
||||
# 3. the alarm truth table (fired / true? and should-have-fired / did it?)
|
||||
# 4. the background household loop's per-round failure summary
|
||||
set -u
|
||||
E=/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17
|
||||
OUT="$E/phase2-morning-after.txt"
|
||||
say(){ echo "$(date -u +%FT%TZ) $*" | tee -a "$OUT"; }
|
||||
H(){ ssh hp "$@" 2>/dev/null | grep -vE "locale|LC_|LANG|perl:|supported and installed|are supported"; }
|
||||
|
||||
say "=== 1. every app, through its own front door ==="
|
||||
H "qm guest exec 336 -- bash -c 'pct exec 9000 -- docker ps --format \"{{.Names}}\t{{.Status}}\"'" | tee -a "$OUT"
|
||||
|
||||
say "=== 1b. the backup page, per tier, read as a first-timer would ==="
|
||||
say "(quoted verbatim into the findings doc; searched with ASCII fragments + controls)"
|
||||
|
||||
say "=== 3. the alarm truth table inputs ==="
|
||||
say "every alarm the hub RECORDED for this customer tonight, with its delivery status."
|
||||
say "'suppressed' and 'never fired' are DIFFERENT and are not allowed to collapse into one."
|
||||
|
||||
say "=== 4. the background household loop ==="
|
||||
if [ -f "$E/household.log" ]; then
|
||||
say "total household operations: $(wc -l < "$E/household.log")"
|
||||
say "failures: $(grep -cE 'FAILED|UNREACHABLE' "$E/household.log")"
|
||||
say "--- failures grouped by 25-minute round window ---"
|
||||
awk '/FAILED|UNREACHABLE/{print substr($1,12,2)":"substr($1,15,1)"0"}' "$E/household.log" | sort | uniq -c | tee -a "$OUT"
|
||||
else
|
||||
say "NO household log — the loop did not run, and that is recorded as a gap, not glossed over."
|
||||
fi
|
||||
say "=== end of the morning-after collection ==="
|
||||
@@ -0,0 +1,3 @@
|
||||
bind submitted at 2026-09-16T20:18:15Z / 22:18 CEST
|
||||
POST bind -> 200
|
||||
page says: Felhom — Doboz összekötése Felhom doboz összekötése Sikeres összekötés. A doboz kb. egy percen belül folytatja a telepítést. Ezt az oldalt bezárhatod — a beállítás a háttérben befejeződik, és a vezérlőpultod hamarosan elérhető lesz. Felhom.eu
|
||||
@@ -0,0 +1,34 @@
|
||||
2026-09-16T20:20:45Z hub-based day-0 watcher started (no SSH needed: the hub sees guests, agent and controller version)
|
||||
2026-09-16T20:20:46Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:21:16Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:21:46Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:22:17Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:22:47Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:23:17Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:23:47Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:24:18Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:24:48Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:25:18Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:25:48Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:26:19Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:26:49Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:27:19Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:27:50Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:28:20Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:28:50Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:29:20Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:29:51Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:30:21Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:30:51Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:31:21Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:31:52Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:32:22Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:32:53Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:33:23Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:33:53Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0
|
||||
2026-09-16T20:34:23Z tester-1-022354 Tester 1 0.131.0 ONLINE 1/1 16% 27% 40% inactive 100% local-lvm Felhom Hub 0.116.0
|
||||
2026-09-16T20:34:23Z GUEST CREATED
|
||||
2026-09-16T20:34:23Z --- controller version the new guest landed on (asked of the customer page) ---
|
||||
Controller 0.245.0
|
||||
Last report: 4 min ago
|
||||
2026-09-16T20:34:24Z day-0 watcher done
|
||||
@@ -0,0 +1,38 @@
|
||||
## NINE DEPLOYS FAILED — and the product said so, correctly, within two minutes
|
||||
## 2026-09-16, box tester-1-022354 / guest 9201, controller 0.245.0
|
||||
|
||||
### What actually happened (root cause, measured)
|
||||
Ten deploys were admitted at 20:30:19, each passing the controller's memory check, which counts
|
||||
COMMITTED memory rather than memory in use:
|
||||
total=4096MB reserved=384MB usable=3712MB — committed_used climbed 2514 -> 3626MB
|
||||
the 11th and 12th (mealie, adventurelog) were REFUSED with the numbers in the message
|
||||
Then the image pulls began, and at **20:34 the hub recorded**:
|
||||
`storage_fill_critical` (critical) — „Host tester-1-022354: storage \"local-lvm\" CRITICALLY full
|
||||
at 100% (threshold 95%) — backups/writes to it will fail; free space immediately"
|
||||
Between 20:32 and 20:33, **nine `app_deploy_failed` (warning) events** were pushed, one per app, each
|
||||
quoting the failing pull: Gokapi, Paperless-ngx, Vaultwarden, Immich, BookStack, Nextcloud, Homebox,
|
||||
Jellyfin, Uptime Kuma. Only **PrivateBin** completed (86 s, `app_deployed` at 20:31:45).
|
||||
|
||||
The disk is thin-provisioned: the VM's 32 GB system disk yields an ~11.8 GB LVM thin pool, over which
|
||||
the guest's 32 GB rootfs and 70 GB data volume are over-subscribed. Ten simultaneous image pulls
|
||||
filled the pool.
|
||||
|
||||
### What this says about the PRODUCT — it behaved, and two of tonight's own concerns are answered
|
||||
* **R-536's `app_deploy_failed` works.** That event shipped this morning precisely so an
|
||||
interrupted install is not silence. Nine interrupted installs produced nine warnings, each
|
||||
naming the app and the failing image, within ~2 minutes of the failure. Before R-536 this was
|
||||
silence, and the customer would have been left with nine cards that never resolved.
|
||||
* **The fill alarm fired at CRITICAL severity** with an actionable sentence, at 95 % threshold.
|
||||
* **The memory guard refused rather than over-committing**, and quoted both numbers it compared.
|
||||
* The box's own per-stack record is honest: `deployed: false` for all nine, `true` only for
|
||||
privatebin. Nothing claims to be installed that is not.
|
||||
|
||||
### What this says about MY HARNESS — two errors, both mine
|
||||
1. **Twelve deploys fired in two seconds is not household behaviour.** A household installs an app,
|
||||
waits for it, then installs another. Firing them in parallel is what drove committed memory to
|
||||
the ceiling and ten image pulls onto one thin pool at once.
|
||||
2. **The drill VM's system disk (32 GB) is too small for a twelve-app household.** That size was
|
||||
copied from a previous drill's VM without checking what that drill actually installed.
|
||||
|
||||
Neither is a product defect and neither is recorded as one. The recovery, and the re-seed done one
|
||||
app at a time, are recorded next.
|
||||
@@ -0,0 +1,602 @@
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:31:42Z containers=4 app.yaml files=10 available_MB=3797
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:32:30Z containers=5 app.yaml files=10 available_MB=3720
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:33:18Z containers=5 app.yaml files=10 available_MB=3818
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:34:05Z containers=5 app.yaml files=10 available_MB=3873
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:34:53Z containers=5 app.yaml files=10 available_MB=3891
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:35:41Z containers=5 app.yaml files=10 available_MB=3890
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:36:29Z containers=5 app.yaml files=10 available_MB=3900
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:37:17Z containers=5 app.yaml files=10 available_MB=3900
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:38:05Z containers=5 app.yaml files=10 available_MB=3899
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:38:53Z containers=5 app.yaml files=10 available_MB=5947
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
lxc-attach: 9201: ../src/lxc/attach.c: get_attach_context: 406 Connection refused - Failed to get init pid
|
||||
lxc-attach: 9201: ../src/lxc/attach.c: lxc_attach: 1474 Connection refused - Failed to get attach context
|
||||
2026-09-16T20:39:41Z containers=0 app.yaml files= available_MB=
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:40:28Z containers=5 app.yaml files=10 available_MB=5933
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:41:16Z containers=5 app.yaml files=10 available_MB=5895
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:42:05Z containers=9 app.yaml files=10 available_MB=4115
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:42:53Z containers=15 app.yaml files=10 available_MB=3704
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:43:42Z containers=17 app.yaml files=10 available_MB=3714
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:44:31Z containers=21 app.yaml files=11 available_MB=3646
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:45:20Z containers=23 app.yaml files=12 available_MB=3131
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:46:09Z containers=25 app.yaml files=12 available_MB=3310
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:46:57Z containers=26 app.yaml files=12 available_MB=4344
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:47:45Z containers=26 app.yaml files=12 available_MB=4292
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:48:33Z containers=26 app.yaml files=12 available_MB=4489
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:49:21Z containers=26 app.yaml files=12 available_MB=4507
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:50:09Z containers=26 app.yaml files=12 available_MB=4503
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:50:57Z containers=26 app.yaml files=12 available_MB=4489
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:51:45Z containers=26 app.yaml files=12 available_MB=4493
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:52:33Z containers=26 app.yaml files=12 available_MB=4464
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:53:21Z containers=26 app.yaml files=12 available_MB=4492
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:54:09Z containers=26 app.yaml files=12 available_MB=4408
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
2026-09-16T20:54:57Z containers=26 app.yaml files=12 available_MB=4479
|
||||
@@ -0,0 +1,7 @@
|
||||
2026-09-16T20:30:12Z pbs_storage=none escrow.pbs_storage_id=none
|
||||
2026-09-16T20:31:13Z pbs_storage=none escrow.pbs_storage_id=none
|
||||
2026-09-16T20:32:14Z pbs_storage=none escrow.pbs_storage_id=none
|
||||
2026-09-16T20:33:15Z pbs_storage=none escrow.pbs_storage_id=none
|
||||
2026-09-16T20:34:15Z pbs_storage=none escrow.pbs_storage_id=none
|
||||
2026-09-16T20:35:16Z pbs_storage=felhom-pbs escrow.pbs_storage_id=felhom-pbs
|
||||
CONVERGED
|
||||
@@ -0,0 +1,60 @@
|
||||
## THE RECOVERY-CODE STEP FAILED ON A FRESH BOX — recorded verbatim, before any retry
|
||||
## 2026-09-16, box `tester-1-022354` / guest 9201, controller 0.245.0, agent 0.131.0
|
||||
|
||||
Context: this is the step controller v0.245.0 added to `VOLUNTEER-first-hour.md` a few hours ago —
|
||||
„A helyreállítási kód (~2 perc) — ezt ne hagyd ki", placed deliberately right after the dashboard
|
||||
password and BEFORE the first app, because until it is done the off-site backup does not run.
|
||||
|
||||
The box was at exactly that point of the guide: installed, bound with no press, claimed, data drive
|
||||
initialised, no apps yet. The escrow reminder bar was on every page, telling the household to do
|
||||
precisely this.
|
||||
|
||||
GET /api/escrow/preflight -> (empty response body)
|
||||
POST /api/escrow/start -> 200 {"job_id":"escrow-1789590499667361664","phase":"running"}
|
||||
GET /api/escrow/status -> claimable:false, claimed:false, and:
|
||||
|
||||
detail: "exit 2: exit status 2 | stderr: selftest=escrow-create requires -storage
|
||||
<pbs-storage-id> (or escrow.pbs_storage…"
|
||||
|
||||
POST /api/escrow/claim -> **409**
|
||||
„A folyamat jelenlegi állapotában a kód nem kérhető le."
|
||||
|
||||
So: the ceremony starts, the agent's `escrow-create` selftest refuses for a missing PBS storage id,
|
||||
and the one-shot claim then correctly declines. The refusal is fail-closed and the wording is honest
|
||||
— nothing pretended to succeed. What is wrong is that the household is TOLD to do this now, on every
|
||||
page, and at this moment it cannot be done.
|
||||
|
||||
Nothing was retried before this file was written, so the state above is the state the box was in.
|
||||
|
||||
## WHY it refused — measured on both sides, not guessed
|
||||
|
||||
**The hub's own Backup & DR panel says it in plain words:**
|
||||
host enrolled (tester-1-022354) done
|
||||
WG tunnel peer registered done
|
||||
descriptor provisioned (namespace tester-1, token felhom@pbs!tester-1) **waiting**
|
||||
„ceremony possible once the descriptor is applied on the box"
|
||||
|
||||
**The box agrees:**
|
||||
pvesm status -> only `local` (dir) and `local-lvm` (lvmthin). **No PBS storage exists yet.**
|
||||
/etc/pve/storage.cfg -> no `pbs:` entry
|
||||
/etc/felhom-agent/agent.json -> top-level keys are
|
||||
[authz, backup, deployment_mode, hub, lan_resolver, local_api, log_level, oob, privileged,
|
||||
proxmox, storage, wg_tunnel]
|
||||
— there is **no `escrow` section at all**, so `escrow.pbs_storage_id` is unset, which is
|
||||
precisely what the agent's selftest complained about.
|
||||
|
||||
**The controller, meanwhile, already has the off-site target:**
|
||||
offbox present=True, enabled=True, host=u629488-sub4.your-storagebox.de, escrow_state=**pending**
|
||||
|
||||
So the chain is: the hub provisioned the DR descriptor automatically at 20:19 (no press), the
|
||||
controller already knows its off-site destination, but the AGENT has not yet applied the descriptor
|
||||
on the box — and the escrow ceremony depends on that. The hub documents the dependency in the very
|
||||
panel that shows it as „waiting".
|
||||
|
||||
## The question this does NOT yet answer, and how it is being measured
|
||||
Whether the box applies the descriptor **by itself**, and how long that takes. That decides
|
||||
everything about severity: a few minutes of convergence makes the new guide step slightly too early
|
||||
in the journey; never converging without an operator press makes it a broken promise on every fresh
|
||||
box. A watcher is now polling the box for `pvesm` gaining a PBS storage and the agent config gaining
|
||||
`escrow.pbs_storage_id`, and the ceremony will be retried when it does. Nothing was pressed, and the
|
||||
box is being left to do it alone.
|
||||
@@ -0,0 +1,62 @@
|
||||
session=64 csrf=64
|
||||
--- preflight ---
|
||||
{"data":{"agent_supported":true,"escrow_state":"pending","items":[{"id":"pbs_storage_id","ok":true,"detail":"felhom-pbs"},{"id":"dr_tier","ok":true,"detail":"DR tier applied"},{"id":"age_binary","ok":true,"detail":"/usr/bin/age"},{"id":"hub_upload","ok":true,"detail":"hub upload target configured"},{"id":"staged_secret","ok":true,"detail":"staged secret present"},{"id":"sudo_grant","ok":true,"deta
|
||||
--- start (re-authenticates with the dashboard password) ---
|
||||
POST /api/escrow/start -> 200
|
||||
{"data":{"job_id":"escrow-1789591140971464235","phase":"running"},"error":"","ok":true}
|
||||
|
||||
--- poll ---
|
||||
{"data":{"claim_expires_in_sec":598,"claimable":true,"claimed":false,"detail":"","entropy_bits":129.24070185585344,"job_id":"escrow-1789591140971464235","key_fingerprint":"6b:ca:5f:3f:ca:0f:e2:3f:fb:2
|
||||
--- claim (ONE-SHOT reveal; value goes to a file, never to stdout) ---
|
||||
POST /api/escrow/claim -> 200
|
||||
recovery code claimed: 83 chars, 10 words — written to a file, NOT printed
|
||||
--- escrow state after the ceremony ---
|
||||
|
||||
## THE ANSWER: the gap is a TIMING gap, and the box closes it BY ITSELF
|
||||
The watcher left the box alone and polled. Nothing was pressed.
|
||||
|
||||
20:30:12Z pbs_storage=none escrow.pbs_storage_id=none
|
||||
20:31:13Z pbs_storage=none escrow.pbs_storage_id=none
|
||||
20:32:14Z pbs_storage=none escrow.pbs_storage_id=none
|
||||
20:33:15Z pbs_storage=none escrow.pbs_storage_id=none
|
||||
20:34:15Z pbs_storage=none escrow.pbs_storage_id=none
|
||||
**20:35:16Z pbs_storage=felhom-pbs escrow.pbs_storage_id=felhom-pbs -> CONVERGED**
|
||||
|
||||
That is **~17 minutes after the bind** (20:18:15Z) and ~6 minutes after the ceremony first refused.
|
||||
The retry then passed every preflight item:
|
||||
pbs_storage_id ok (felhom-pbs) · dr_tier ok (DR tier applied) · age_binary ok (/usr/bin/age)
|
||||
· hub_upload ok · staged_secret ok · sudo_grant ok · agent_supported true · escrow_state pending
|
||||
|
||||
POST /api/escrow/start -> 200, job escrow-1789591140971464235, phase running
|
||||
GET /api/escrow/status -> claimable:true, entropy_bits 129.2, key fingerprint 6b:ca:5f:3f:…
|
||||
POST /api/escrow/claim -> **200** — the recovery code, 83 characters / 10 words, revealed ONCE
|
||||
(written to a 0600 file; not printed, not committed — the household writes it on paper)
|
||||
|
||||
**So the product is not broken here, and the fix shipped tonight is not wrong — the GUIDE's timing
|
||||
is.** `VOLUNTEER-first-hour.md` §6 (written a few hours ago) places the recovery code immediately
|
||||
after the dashboard password and before the first app. On a fresh box that moment is inside the
|
||||
~17-minute window where the agent has not yet applied the DR descriptor, so a volunteer following the
|
||||
guide literally meets „exit 2: selftest=escrow-create requires -storage <pbs-storage-id>" and a 409,
|
||||
with nothing on the page telling them to simply wait a quarter of an hour.
|
||||
|
||||
The escrow reminder bar (R-543, also shipped tonight) makes this sharper rather than softer: it is on
|
||||
every page urging the household to do the very thing that cannot yet be done.
|
||||
|
||||
## The state flipped: escrow_state = **escrowed**
|
||||
Read from the box's own settings after the ceremony:
|
||||
enabled=True host=u629488-sub4.your-storagebox.de
|
||||
**escrow_state=escrowed**
|
||||
last_run=None last_status=None snapshot_count=None (never run yet — honest, not "0")
|
||||
|
||||
„Claimed" and „escrowed" are different facts: the claim is the household seeing the code once, the
|
||||
flip to `escrowed` happens when the hub's ACK confirms `sha256(local repo_password)` matches the
|
||||
stored escrow. Both happened. The off-site tier is therefore ARMED for the first time on this box,
|
||||
and the escrow reminder bar shipped tonight should now be gone from every page — which the first
|
||||
round will read back rather than assume.
|
||||
|
||||
## An alarm caused by MY recovery, flagged so it is never scored as a round's alarm
|
||||
Sep 16 20:39 **error storage_disconnected** „Meghajtó váratlanul leválasztva: Adatlemez"
|
||||
That is the guest restart I performed to clear the read-only wedge. It is a TRUE alarm — the drive
|
||||
really did go away for those seconds — but it belongs to my repair, not to any accident in the
|
||||
schedule. Round 11's drawn accident is `drive pulled 20min`, and when that round is scored this
|
||||
20:39 event must not be mistaken for it.
|
||||
@@ -0,0 +1,9 @@
|
||||
## events baseline, captured 2026-09-16T20:15:23Z, BEFORE round 1
|
||||
## everything above this line in later dumps belongs to yesterday's box, not tonight's
|
||||
| Time | Severity | Type | Message | Source
|
||||
| Sep 16 18:47 | error | node_down | No report received for 1h | hub
|
||||
| Sep 16 18:17 | info | selfbind_link_sent | Self-bind link e-mailed (host delete) | hub
|
||||
| Sep 16 18:17 | warning | node_stale | No report received for 30m | hub
|
||||
| Sep 16 18:16 | warning | host_stale | Host tester-1-33b6a9: no report for 30m | hub
|
||||
| Sep 16 17:25 | info | pbsdr_adopted | PBS DR token adopted by operator re-issue (endpoint held a token, hub had no descriptor) | hub
|
||||
| Sep 16 17:16 | info | app_deployed | Alkalmazás telepítve: Nextcloud | controller
|
||||
@@ -0,0 +1,66 @@
|
||||
2026-09-16T20:11:58Z first-boot watcher started (box 192.168.0.115, VM 336)
|
||||
2026-09-16T20:11:58Z PVE port 8006 answers
|
||||
2026-09-16T20:14:57Z watcher restarted (the previous one killed ITSELF: pkill -f matched its own command line)
|
||||
2026-09-16T20:14:58Z bootstrap=activating agent=none guests=0
|
||||
2026-09-16T20:15:20Z bootstrap=activating agent=none guests=0
|
||||
2026-09-16T20:15:42Z bootstrap=activating agent=none guests=0
|
||||
2026-09-16T20:16:04Z bootstrap=activating agent=none guests=0
|
||||
2026-09-16T20:16:25Z bootstrap=activating agent=none guests=0
|
||||
2026-09-16T20:16:47Z bootstrap=activating agent=none guests=0
|
||||
2026-09-16T20:17:09Z bootstrap=activating agent=none guests=0
|
||||
2026-09-16T20:17:30Z bootstrap=activating agent=none guests=0
|
||||
2026-09-16T20:17:52Z bootstrap=activating agent=none guests=0
|
||||
2026-09-16T20:18:14Z bootstrap=activating agent=none guests=0
|
||||
2026-09-16T20:18:36Z bootstrap=activating agent=none guests=0
|
||||
2026-09-16T20:18:58Z bootstrap=activating agent=none guests=0
|
||||
2026-09-16T20:19:27Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:19:56Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:20:25Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:20:54Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:21:23Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:21:50Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:22:19Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:22:48Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:23:17Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:23:46Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:24:15Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:24:44Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:25:12Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:25:41Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:26:10Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:26:38Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:27:07Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:27:36Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:28:05Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:28:32Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:29:01Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:29:30Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:29:59Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:30:28Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:30:58Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:31:27Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:31:56Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:32:25Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:32:54Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:33:23Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:33:52Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:34:21Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:34:50Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:35:19Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:35:48Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:36:17Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:36:46Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:37:15Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:37:44Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:38:13Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:38:42Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:39:11Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:39:40Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:40:09Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:40:38Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:41:07Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:41:36Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:42:05Z bootstrap= agent=none guests=0
|
||||
2026-09-16T20:42:25Z --- bootstrap journal, last 40 ---
|
||||
2026-09-16T20:42:28Z --- what it landed on ---
|
||||
2026-09-16T20:42:31Z watcher done
|
||||
@@ -0,0 +1,17 @@
|
||||
## 2026-09-16T19:52:02Z golden 0.245.0 — bake in the drill VM (RUNBOOK-manual-build §4.0/§4.1)
|
||||
reverted to virgin (this also proves no qemu holds the qcow2)
|
||||
cold boot started
|
||||
ssh up: pve-manager/9.2.2/b9984c6d90a4bd80 (running kernel: 7.0.2-6-pve)
|
||||
template: debian-13-standard_13.6-1_amd64.tar.zst
|
||||
template downloaded
|
||||
script + token landed, non-empty, executable
|
||||
bake launched 2026-09-16T19:52:47Z (GITEA_USER=admin — the 'kisfenyo' namespace error cost a bake)
|
||||
token leak check on the unit (needle proven non-empty, so the grep cannot match everything): 0
|
||||
unit state: inactive at 2026-09-16T19:58:15Z
|
||||
log copied off the machine FIRST: 325 lines
|
||||
markers: overlay2=1 mountpoints=2 upload=1 FATAL=0 skipped=0
|
||||
GOLDEN_VERSION=0.245.0
|
||||
GOLDEN_SHA256=7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626
|
||||
REGISTRY CHECK (the outcome, not the attempt): golden 0.245.0 -> http=200
|
||||
PUBLISHED; guest destroyed, qemu exited, disk reverted to virgin
|
||||
## bake finished 2026-09-16T19:58:52Z
|
||||
@@ -0,0 +1,21 @@
|
||||
install pressed at 2026-09-16T20:07:11Z
|
||||
2026-09-16T20:07:33Z watcher started; completion is judged from the DISK (blocks actually used), never from the screen
|
||||
2026-09-16T20:07:34Z system-disk blocks used: 3233 MiB (stable checks: 0)
|
||||
2026-09-16T20:08:04Z system-disk blocks used: 5375 MiB (stable checks: 0)
|
||||
2026-09-16T20:08:35Z system-disk blocks used: 6762 MiB (stable checks: 0)
|
||||
2026-09-16T20:09:06Z system-disk blocks used: 6843 MiB (stable checks: 0)
|
||||
2026-09-16T20:09:36Z system-disk blocks used: 6942 MiB (stable checks: 0)
|
||||
2026-09-16T20:10:07Z system-disk blocks used: 6942 MiB (stable checks: 1)
|
||||
2026-09-16T20:10:37Z system-disk blocks used: 6942 MiB (stable checks: 2)
|
||||
2026-09-16T20:11:08Z system-disk blocks used: 6942 MiB (stable checks: 3)
|
||||
2026-09-16T20:11:08Z INSTALL COMPLETE (disk stopped growing at 6942 MiB)
|
||||
2026-09-16T20:11:08Z applying the post-install fix: stop, detach the CD, boot order scsi0, start from disk
|
||||
disk usage before: 6942 MiB
|
||||
update VM 336: -delete ide2
|
||||
update VM 336: -boot order=scsi0
|
||||
boot: order=scsi0
|
||||
scsi0: nvme-scratch:336/vm-336-disk-1.raw,size=32G
|
||||
scsi1: nvme-scratch:336/vm-336-disk-0.raw,size=100G
|
||||
scsihw: virtio-scsi-single
|
||||
started from disk at 2026-09-16T20:11:19Z
|
||||
2026-09-16T20:11:19Z post-install fix done
|
||||
@@ -0,0 +1,362 @@
|
||||
## CHAOS NIGHT — Phase 0 notes (2026-09-16 evening, CEST)
|
||||
|
||||
### Baselines, re-verified live against Gitea at 21:49 CEST (not copied from the brief)
|
||||
felhom-controller 714d5bce0920 v0.245.0 (MinAgent 0.131.0) — clean, in sync
|
||||
felhom-agent e98b857684f4 v0.131.0 — clean, in sync
|
||||
felhom.eu d124c77e176d hub v0.116.0, ISO 1.28.0 live — clean, in sync
|
||||
app-catalog 94bc5febaca2 — clean, in sync
|
||||
Register: highest R-545, 212 open. Golden waiver valid to 2026-09-27.
|
||||
|
||||
### A claim in the brief, CHECKED rather than inherited
|
||||
The brief says the customer `tester-1` has "no host". CONFIRMED: /hosts lists exactly three hosts —
|
||||
demo-felhom-8363b5, demo-hp-bb76ea, drill-r50-0a4f9a. There is no tester-1 host record. The
|
||||
customers list showing „tester-1 … DOWN … 0.244.0" is the customer's LAST KNOWN state, not a live
|
||||
box; reading that row as a host record would have been the mistake.
|
||||
|
||||
### 0.1 — golden 0.245.0, baked and published
|
||||
Launched 19:52:47Z as a transient unit in the drill VM; finished 19:58:15Z.
|
||||
markers: overlay2=1 mountpoints=2 upload=1 FATAL=0 publish-skipped=0
|
||||
GOLDEN_VERSION=0.245.0
|
||||
GOLDEN_SHA256=7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626
|
||||
REGISTRY CHECK (the outcome, not the attempt): golden 0.245.0 -> http=200
|
||||
then: guest 9100 destroyed, token shredded, qemu exited, disk reverted to virgin.
|
||||
|
||||
All three of 2026-09-16's bake failures were guarded against and none recurred: `scp -P` (not the
|
||||
ssh `-p`), `chmod 0700` on the script, and `GITEA_USER=admin` (not the first credentials line).
|
||||
The token-leak check ran with a needle PROVEN non-empty first, because `grep -F ""` matches every
|
||||
line and an instrument that reports a hit on an empty needle is not a measurement.
|
||||
|
||||
HONEST NOTE ON THE EXIT CODE: the wrapper script exited **144** while every measured outcome was
|
||||
good. That is why teardown was gated on the REGISTRY answering 200 and not on an exit code — this
|
||||
repo's own "exit codes that lie" class. The artifact is published and verified independently.
|
||||
|
||||
### 0.1b — vouched in the hub
|
||||
POST /configuration/artifacts -> 303, then READ BACK from the page (the outcome, not the POST code):
|
||||
golden currently vouched: 0.245.0
|
||||
golden sha now: 7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626
|
||||
agent 0.131.0 and the wrapper sha left exactly as they were.
|
||||
|
||||
DECISION, and the reason, because silence reads as agreement: the GLOBAL controller floor was left at
|
||||
0.244.0 and NOT raised to 0.245.0. The new box installs golden 0.245.0, which already carries
|
||||
controller 0.245.0, so the floor is not needed to deliver anything tonight; raising it would have
|
||||
pushed a controller update onto demo-felhom, a box that is not part of this drill, at 22:00 at night.
|
||||
The brief asked for a bake and a vouch, not a floor raise.
|
||||
|
||||
### 0.2 — the box
|
||||
demo-hp (Tier 0). VM 336 „tester1-chaos-night": 8 GiB, 4 cores, cpu host, virtio-scsi-single,
|
||||
scsi0 = 32 G system disk, scsi1 = 100 G data disk, both on `nvme-scratch` (dir storage, path
|
||||
/mnt/hdd_1, is_mountpoint yes — the NVMe at its ROOT, per target-selection.md). CD-ROM is the
|
||||
PUBLISHED felhom-installer-1.28.0-pve9.2-1.iso. **The boot order was set in its own `qm set`** —
|
||||
combining it with the disk call silently yields `order=net0;ide2` and the VM netboots.
|
||||
boot: order=ide2;scsi0 (verified by reading `qm config 336` back)
|
||||
The VM took DHCP 192.168.0.115 and booted the installer's graphical entry.
|
||||
|
||||
### Console driving — measured, not assumed
|
||||
* Enter on the EULA page: advances.
|
||||
* Enter on the Location page: lands IN the Country text field and does NOT press Next (the
|
||||
documented GTK trap).
|
||||
* **Alt+N works as the „Next" mnemonic** and is what actually drives this installer. Recorded
|
||||
because the previous drill switched to the text-mode entry to avoid exactly this problem.
|
||||
* The VM has a QEMU HID Tablet (absolute), so `mouse_move <x> <y>` takes absolute coordinates and
|
||||
a real click is available as a fallback. `info mice` says so — checked, not assumed.
|
||||
* The installer REFUSES the prefilled `mail@example.invalid` with „Please enter a valid Email
|
||||
address" and simply does not advance. The dialog is above the fold, which is why the cropped
|
||||
strip looked like „nothing happened" — the full screen showed the reason.
|
||||
* Keyboard layout is Hungarian (QWERTZ): `y` and `z` are swapped and `@` is AltGr+V. The drill root
|
||||
password was generated from a-x plus digits so it is identical under either layout, and it is
|
||||
stored out-of-band (0600, scratchpad) — never in a committed file.
|
||||
|
||||
### Console driving on a Hungarian-layout installer — MEASURED, and it cost four round trips
|
||||
These are facts about driving a PVE 9.2 graphical installer headlessly through `qm monitor`, and
|
||||
every one of them was measured on this box tonight rather than recalled:
|
||||
|
||||
* `sendkey alt-n` is the „Next" mnemonic and is what actually advances the installer. Plain `ret`
|
||||
advances ONLY the EULA page; on every later page it lands inside a text entry and does nothing.
|
||||
* **`sendkey altgr-v` does NOTHING.** On a Hungarian layout `@` is AltGr+V, and the QEMU key name
|
||||
that works is **`alt_r`** — `sendkey alt_r-v` types the `@`. The failure is silent: the character
|
||||
is simply absent, so `admin@felhom.eu` became `adminfelhom.eu` and the installer then refused the
|
||||
page with „Please enter a valid Email address" — a refusal that looks exactly like „nothing
|
||||
happened" if you only crop the bottom strip of the screen.
|
||||
* `mouse_move <x> <y>` did not move the pointer even though `info mice` reports a QEMU HID Tablet
|
||||
(absolute) as the active device. Clicking was abandoned; keyboard navigation is the reliable path.
|
||||
* **Focus is found by MEASURING it, not by counting tabs.** A 20-line script samples the blue border
|
||||
of each entry box in the screendump and prints which one is focused; Tab is then pressed until
|
||||
the wanted field reports focus. Counting keystrokes blind is how a password ends up in the wrong
|
||||
field. (`blueness = mean(B-R)` over the box's top border: focused +58, unfocused 0.)
|
||||
* The installer REFUSES the prefilled `mail@example.invalid`.
|
||||
|
||||
### Fence check at this point
|
||||
demo-hp: /mnt/hdd_1 has 876 G free; VM 336's two raw disks are sparse (132 G apparent, 12 K actual).
|
||||
`/` on demo-hp is at 77 % and was deliberately NOT used for VM disks. Guests 9201 and 9202 untouched
|
||||
and running (they are the standing demo boxes, and 9202 is the scratch guest the morning restore
|
||||
will use). No `local-lvm`, no prune, `drill-r50` not touched.
|
||||
|
||||
### A second brief claim, CHECKED rather than inherited: „the automatic mail is waiting in the mailbox"
|
||||
CONFIRMED. The mailbox holds three „[Felhom] Kösd össze a Felhom dobozodat" messages to
|
||||
`tester1@felhom.eu` from `monitoring@felhom.eu`, the newest at **2026-09-16T18:17:46Z** — the same
|
||||
second as yesterday's host delete, which is R-509's automatic trigger firing. So the box installed
|
||||
tonight should be bindable with **no operator press**, and the pre-declared intervention O1 (the
|
||||
„Send self-bind link" button) should not be needed. Whether it IS needed is a measurement of this
|
||||
night, not an assumption: if the waiting link turns out to be superseded or refused, that is a
|
||||
finding and the press becomes O1.
|
||||
|
||||
The other mail in that mailbox worth noting, because it is the same customer's history and could
|
||||
confuse a reader of this evidence: two „Új beállító kód — újratelepült a szervered" messages
|
||||
(10:01:57Z and 16:00:58Z) and one „Beállító kód a jelszavad visszaállításához" (11:12:06Z), all from
|
||||
2026-09-16 — those belong to yesterday's drills, not to tonight's box.
|
||||
|
||||
NOT recorded here, deliberately: the bind link itself. It carries a one-time token, and a one-time
|
||||
secret does not go into a committed file (nor was the mail body fetched into the session transcript
|
||||
for the same reason — the link will be taken from the hub at bind time instead).
|
||||
|
||||
### The one catalog bump — what it is, and why the ORDER matters
|
||||
Round 7 drew `update nextcloud`, so the bump has to be on the **nextcloud** template, not on whatever
|
||||
small app I would have picked. The brief asked for "one drill bump of one SMALL app"; the draw
|
||||
overrides the choice of app, so the compromise is to bump the template's small sidecar pin rather
|
||||
than the big application image:
|
||||
|
||||
templates/nextcloud/docker-compose.yml:103 redis:7-alpine -> redis:7.4-alpine
|
||||
|
||||
`redis:7.4-alpine` was CHECKED to exist on Docker Hub before being written down — inventing a tag
|
||||
would have made round 7 fail for the wrong reason, and a round that fails for the wrong reason
|
||||
measures nothing. `catalog_since` moves with it, per this repo's own rule that any commit changing an
|
||||
`image:` line must.
|
||||
|
||||
**The bump is NOT pushed yet, and the order is the point:** nextcloud must be DEPLOYED from the
|
||||
current catalog first, or the box installs 7.4-alpine immediately and round 7 has no update to
|
||||
apply. Sequence: seed the twelve apps → then push the bump → the box's git-sync picks it up within
|
||||
15 min (or the „Sablonok frissítése" button) → round 7 at 02:00 has a real update waiting.
|
||||
|
||||
Fence note: the catalog is SHARED with the demo boxes. A bump offers an update; it never applies one
|
||||
— the guarded update needs a press — so the demo boxes' standing apps are not disturbed by this.
|
||||
|
||||
### Deploy fields, read from the templates rather than guessed
|
||||
All twelve templates live under `templates/<app>/`, not at the repo root (my first lookup used the
|
||||
wrong layout and returned "NO .felhom.yml" twelve times — recorded because the wrong answer was
|
||||
uniform and therefore looked authoritative). Only SUBDOMAIN and HDD_PATH are marked `required`;
|
||||
every secret field is optional in metadata but the server refuses a deploy without them (the
|
||||
metadata-vs-server contradiction measured 2026-09-16), so the seeding script supplies generated
|
||||
values for all of them and writes them to a 0600 file that is never echoed.
|
||||
Apps needing the data drive: nextcloud, immich, paperless-ngx, jellyfin.
|
||||
|
||||
### 0.2b — the install itself
|
||||
Install pressed 20:07:11Z. **Completion was judged from the DISK, not from the screen** — a PVE
|
||||
install with "Automatically reboot" ticked re-enters the installer, so a finished install and a stuck
|
||||
one look identical on screen. The system disk's blocks-used grew 3233 → 6942 MiB and then held
|
||||
steady across three consecutive 30-second checks:
|
||||
20:09:36Z 6942 MiB (stable 0) · 20:10:07Z 6942 (1) · 20:10:37Z 6942 (2) · 20:11:08Z 6942 (3)
|
||||
-> INSTALL COMPLETE 20:11:08Z
|
||||
|
||||
Then the documented post-install fix, because a RUNNING guest keeps the boot order QEMU started
|
||||
with — editing the config mid-run is not enough:
|
||||
qm stop 336 · qm set 336 --delete ide2 · qm set 336 --boot order=scsi0 (its own call) · qm start
|
||||
read back: `boot: order=scsi0`, no ide2, both disks intact
|
||||
started from disk 20:11:19Z
|
||||
|
||||
The summary page was checked before Install was pressed, and the line that mattered was
|
||||
**„Disk(s): /dev/sda"** — the 32 G system disk alone. The 100 G data disk was NOT offered to the
|
||||
installer and was not touched. That is the check that makes pressing Install safe on a box with a
|
||||
second disk, and it is the one a two-disk filter has silently got wrong elsewhere in this project.
|
||||
|
||||
### How this box is driven, and how the alarms are read — method facts, each one measured
|
||||
* **`qm guest exec` is UNUSABLE on this box.** The VM was created with `--agent 1`, but a plain PVE
|
||||
install does not run `qemu-guest-agent`, and `qm agent 336 ping` answers nothing. My first
|
||||
first-boot watcher was written against `qm guest exec` and therefore sat silent while reporting
|
||||
nothing — it looked like a box that would not boot, and it was an instrument pointed at nothing.
|
||||
* **The access path is SSH to the box** (`root@192.168.0.115`, the drill password via `sshpass -e`
|
||||
from a 0600 file, never on a command line). Confirmed with `LOGIN_OK`, hostname `chaosnight`,
|
||||
`pve-manager/9.2.2`.
|
||||
* **The alarm instrument is the hub's own `events-table`** on `/customers/tester-1` (note: the
|
||||
customer page is `/customers/<id>`, NOT `/configs/<id>` — the latter redirects). The tabs are
|
||||
client-side, so `?tab=events` returns the same document; the events must be pulled out of the
|
||||
table by its container id. The hub pod has **no `sqlite3`** and `hub.db` is 298 MB with a live
|
||||
`-wal`, so copying the DB is both heavy and stale-prone; the page is the correct instrument.
|
||||
* **Events baseline, captured before round 1:** the newest pre-drill event is
|
||||
`Sep 16 18:47 error node_down` (yesterday's box). Anything newer belongs to tonight. Without this
|
||||
marker, yesterday's `node_stale`/`node_down`/`selfbind_link_sent` rows would be scored as
|
||||
tonight's alarms.
|
||||
|
||||
### My own mistake, recorded because it cost two watchers
|
||||
`pkill -f "first-boot watcher"` **matched its own command line** and killed the very background job
|
||||
it was meant to clear, twice, each exiting 144. Same class as the documented `pgrep -f
|
||||
qemu-system-x86_64` self-match. The replacement watcher does no pkill at all.
|
||||
|
||||
### Fence note on a secret
|
||||
While following redirects to find the customer page, `curl -w '%{url_effective}'` printed the hub
|
||||
password back to me inside the resolved URL. It is not in any file written here, and that format
|
||||
option is not used again. Recorded rather than quietly dropped, because the next person will hit the
|
||||
same flag.
|
||||
|
||||
### 0.2c — the bind, as a volunteer does it: ZERO operator presses
|
||||
The box registered itself as an unclaimed appliance and then sat polling every 30 s — correctly, and
|
||||
visibly: „not bound yet — polling every 30s until the operator or a customer self-bind lands". No
|
||||
agent and no guest exist until the bind happens, so a reader who expected the box to finish
|
||||
installing by itself would have mis-read a waiting box as a stuck one.
|
||||
|
||||
The volunteer's own path was taken, end to end:
|
||||
* the link came from the **waiting mail** (minted 2026-09-16T18:17:46Z by yesterday's host delete,
|
||||
R-509's automatic trigger) — not from the operator's „Send self-bind link" button;
|
||||
* the **pairing code** was read off the box's own console: `4SY-4TX`;
|
||||
* the **„Tulajdonosi jelmondat"** (5 words) came from the hub's customer record, which is where the
|
||||
operator hands it from — it is never e-mailed, by design.
|
||||
|
||||
POST /bind/<token> -> 200 at 2026-09-16T20:18:15Z
|
||||
„Sikeres összekötés. A doboz kb. egy percen belül folytatja a telepítést. Ezt az oldalt
|
||||
bezárhatod — a beállítás a háttérben befejeződik, és a vezérlőpultod hamarosan elérhető lesz."
|
||||
|
||||
**O1 (pre-declared) was NOT used and is not counted.** The brief allowed one press of „Send self-bind
|
||||
link" if no mail was waiting; a mail WAS waiting and it worked, so the bind cost zero interventions.
|
||||
|
||||
Secrets discipline for this step: the bind token, the passphrase and the box's root password each
|
||||
live in a 0600 scratchpad file and were passed to `curl --data-urlencode name@file`, so no value
|
||||
reached a command line. None of the three is written into this evidence, and the passphrase's only
|
||||
description here is its shape (5 words, 35 characters).
|
||||
|
||||
### 0.2d — the box rotates its own root password at day-0, and my access died with it
|
||||
At 20:18:58Z my SSH to the box still worked; at 20:19:04Z it answered „Permission denied", and the
|
||||
background watcher lost access in the same window (its 20:19:27Z line came back empty). The bind at
|
||||
20:18:15Z had started the day-0 install.
|
||||
|
||||
**This is the design, not a defect, and it was confirmed in source rather than guessed:**
|
||||
`scripts/felhom-host-install.sh` (≈2079-2104) generates a strong `root@pam` password with `openssl
|
||||
rand`, sets it through `chpasswd` on **stdin** (no argv, no log), and vaults it to the hub with
|
||||
`PUT /api/v1/hosts/<id>/recovery-credential` over the enroll-authenticated channel. The password is
|
||||
never logged, printed, or written to a file anywhere on the box. So the installer-time password I
|
||||
typed into the Proxmox installer is dead by intent the moment day-0 runs.
|
||||
|
||||
**What this changes for tonight:** every accident that acts INSIDE the box (docker restart, tunnel
|
||||
kill, filling the system disk) needs the hub-vaulted break-glass credential
|
||||
(`/hosts/<host-id>/reveal-recovery-credential`), not the install password. Discovered at 22:19 CEST,
|
||||
before the rounds began, rather than at 01:10 in the middle of round 5 — which is the only reason it
|
||||
is a method note here and not an intervention later.
|
||||
|
||||
**How it was diagnosed honestly:** my first instinct was that I had broken my own environment. That
|
||||
was ruled out first — the password file was still 21 bytes, the variable still 20 characters, and the
|
||||
same credential had worked six seconds earlier. Only then was the box's own behaviour blamed, and
|
||||
only after the installer source confirmed the mechanism.
|
||||
|
||||
### 0.2e — THE F-14 PATH, MEASURED LIVE FOR THE FIRST TIME
|
||||
The brief named this as a claim that had never been measured: „the WG hook provisions by itself after
|
||||
an acknowledged delete". Tonight it ran, unprompted, and the hub recorded it:
|
||||
|
||||
Sep 16 20:18 appliance_bound (customer_selfbind)
|
||||
„Az ügyfél saját maga kötötte össze az új eszközt (bare-metal telepítés); a hozzáférést a
|
||||
doboz a következő lekérdezéskor megkapja."
|
||||
Sep 16 20:18 appliance_credential_delivered
|
||||
„Új eszköz (bare-metal telepítés) megkapta a hozzáférést és megkezdi a beállítást."
|
||||
Sep 16 20:18 claim_reissued_reenroll
|
||||
„Új beállító kódot küldtünk a szerver újratelepítése után (6. generáció) az ügyfél címére."
|
||||
Sep 16 20:19 offsite_reissued
|
||||
„Az offsite (házon kívüli) mentési hozzáférést újra kiadtuk — az új egyszeri jelszót a vezérlő
|
||||
a következő frissítéskor átveszi."
|
||||
Sep 16 20:19 **pbsdr_auto_reissue**
|
||||
„Previous key destroyed (acknowledged deletion) — credentials re-issued automatically."
|
||||
|
||||
That last line is the one that matters. Yesterday's box was removed through the **acknowledged**
|
||||
delete flow, and tonight's box therefore got its off-site and PBS-DR credentials **with no operator
|
||||
press at all** — exactly what the F-14 ruling of 2026-07-13 says should happen on that path, and the
|
||||
half of that ruling nobody had yet watched happen.
|
||||
|
||||
**Consequence for this night's intervention count:** BOTH pre-declared presses are unnecessary.
|
||||
O1 („Send self-bind link") was not needed because the automatic mail was waiting; O2 („Re-issue PBS
|
||||
credentials") was not needed because the acknowledged-delete path re-issued by itself.
|
||||
**Interventions so far: 0.**
|
||||
|
||||
Host enrolled as **`tester-1-022354`**, agent **0.131.0**, ONLINE, guests 0/0 at 20:20Z — the
|
||||
customer guest is still being created from the golden.
|
||||
|
||||
### 0.2f — day-0 delivered tonight's golden, with no hand upgrade
|
||||
guest: 9201 „tester-1", running, created by the bootstrap from the golden
|
||||
controller image: gitea.dooplex.hu/admin/felhom-controller:**0.245.0** — „Up … (healthy)"
|
||||
agent: felhom-agent **0.131.0**
|
||||
host: `tester-1-022354`, ONLINE in the hub
|
||||
|
||||
The golden baked at 19:58Z tonight (sha 7a08aa1a…) is what this box installed. Nothing was upgraded
|
||||
by hand, and the controller the customer will use is the release this drill is validating. That is
|
||||
the delivery half of the chain: bake -> vouch -> a fresh box lands on it.
|
||||
|
||||
both disks present to the box: `sda` 32 G (system, PVE + LVM) and **`sdb` 100 G** (the data disk,
|
||||
still unformatted — the household's drive, initialised through the storage page in the next step).
|
||||
|
||||
MY OWN MEASUREMENT ERROR, recorded: the first `lsblk` was piped through `head -12` and stopped one
|
||||
line short of `sdb`. For a minute the box looked like it had NO data disk — a wrong answer produced
|
||||
entirely by my own truncation, not by the box. Re-read without the pipe, `sdb 100G` is plainly there
|
||||
and `qm config 336` still shows `scsi1` attached. An instrument that can cut off its own answer is
|
||||
not a measurement.
|
||||
|
||||
### the dashboard setup code
|
||||
The hub mailed a fresh „Új beállító kód — újratelepült a szervered" at **20:18:56Z** (the 6th
|
||||
generation for this customer), 72 hours valid, delivered to `tester1@felhom.eu`. That is the code the
|
||||
volunteer types on „A szerver beállítása" to set their own dashboard password — and it arrived by
|
||||
itself, as part of the same automatic re-enrolment that needed no operator press.
|
||||
|
||||
### 0.2g — the dashboard claimed, by the volunteer, with the mailed code
|
||||
The code from the 20:18:56Z mail („ősrégen-újraért-címbetű", 3 words, 72 h) was typed into
|
||||
„A szerver beállítása" together with a 20-character password the household chooses.
|
||||
|
||||
POST /claim -> 302
|
||||
POST /login (new pw) -> 302, and a `felhom_session` cookie was issued
|
||||
|
||||
**The second line is the proof; the first is only an attempt.** This repo has a standing trap that
|
||||
an HTTP 200 (or a redirect) can be a refusal — the claim page re-rendering itself looks exactly like
|
||||
success from the status code alone. The claim is called successful here because the password it set
|
||||
then opened a session, which is the consequence a customer actually cares about.
|
||||
|
||||
Both values went in as FILES (`--data-urlencode name@file`, 0600, pushed with `pct push`), so
|
||||
neither the setup code nor the new password ever reached a command line on the host or in the guest.
|
||||
|
||||
### 0.2h — tonight's release, seen working on a box that installed itself
|
||||
The storage page of this fresh box carries `<form method="POST" action="/backup/escrow/banner/dismiss">`
|
||||
— the **R-543 escrow reminder bar shipped in controller v0.245.0 a few hours ago**, rendering on a
|
||||
box nobody had touched. It is there because the off-site tier was re-issued automatically at 20:19Z
|
||||
and its escrow is not complete yet, which is precisely the state the bar exists for.
|
||||
|
||||
This is the first time that fix has been seen on a box that was not set up for the purpose of
|
||||
testing it: the box installed itself from the published ISO, landed on tonight's golden, got its
|
||||
off-site credentials with no press, and is now telling the household — on every page — that the
|
||||
remote backup is paused until they create their recovery code. Creating it is the next step of the
|
||||
guide, and of this drill.
|
||||
|
||||
### the data drive, as the box offers it
|
||||
`/api/disks/candidates` reports exactly one initialisable device:
|
||||
/dev/sdb — 107 374 182 400 B (100 GB), QEMU HARDDISK, data_bearing=false, mountable=false
|
||||
and separately the guest's own system volume as already mounted at /mnt/sys_drive. The empty
|
||||
100 GB disk is the household's drive and the only thing offered for initialisation — the
|
||||
data_bearing=false flag is the guard that keeps a drive with someone's files on it out of this list.
|
||||
|
||||
### Schedule: Phase 0 ran long, and the rounds shift with it — recorded, not quietly re-timed
|
||||
The drawn schedule puts round 1 at 23:30 CEST. Phase 0 will not be finished by then: the box was
|
||||
installed, bound, claimed and landed on tonight's golden without trouble, but working out how the
|
||||
storage wizard actually initialises a disk took several rounds of discovery, because the wizard
|
||||
submits through JavaScript (`POST /api/storage/init`) rather than a form, and I refused to guess the
|
||||
endpoint after a guessed path cost a 403 and a wrong diagnosis in an earlier drill.
|
||||
|
||||
**What shifts and what does not.** The SCHEDULE — which action, on which app, under which accident,
|
||||
in which order — is unchanged; it was drawn from the seed before anything ran and is fixed. Only the
|
||||
wall-clock start moves, and the ~25-minute spacing is kept from the new start. The brief allows a
|
||||
round to wait provided the wait is recorded; this is that record. The 05:00 stop rule is unchanged,
|
||||
so a late start means the night may reach fewer than twelve rounds, and the morning verdict will say
|
||||
how many actually ran rather than implying all twelve did.
|
||||
|
||||
### 0.2i — the household's drive, initialised through the wizard's own endpoint
|
||||
The wizard submits by JavaScript, not by a form: `POST /api/storage/init` with
|
||||
`{device, fstype, mount_name, label, set_default, confirmed, durable_id}` and the CSRF meta token,
|
||||
then polls `GET /api/storage/init/status`. Both were read off the live page and confirmed in
|
||||
`internal/web/storage_handlers.go:362` before anything was sent.
|
||||
|
||||
POST /api/storage/init -> 200 {"phase":"formatting","started":true}
|
||||
GET /api/storage/init/status -> phase **done**, started 20:27:21.79Z, updated 20:27:23.95Z
|
||||
where=/mnt/felhom-drives/hdd_1, error="" reason=""
|
||||
read back: /dev/sdb is ext4, durable_id `uuid:8f59ed90-c0e9-4e20-8584-d18af823c605`
|
||||
`df`: /dev/sdb 98G, 2.1M used, 93G free, mounted on /mnt/felhom-drives/hdd_1
|
||||
storage page: „hdd_1" labelled „Adatlemez", set as default
|
||||
|
||||
**The POST returning 200 is not the result** — it only says the job started. The format runs as a
|
||||
background job precisely so a closed tab cannot abort it, so the phase poll is what says it worked,
|
||||
and the `df` line is what says the household can use it.
|
||||
|
||||
**A known row met tonight, recorded rather than re-filed: R-542.** After the drive is formatted,
|
||||
registered, mounted and made default, `/api/disks/candidates` STILL lists `/dev/sdb` under
|
||||
`initialize` — now with `data_bearing:true, mountable:true` and its durable id. The same endpoint
|
||||
also lists it under `attach`. That is exactly the behaviour R-542 describes (a registered, in-use
|
||||
drive still offered under „initialize"), seen again on a fresh box.
|
||||
@@ -0,0 +1,126 @@
|
||||
--- the new disk as the box sees it ---
|
||||
sda 32G disk
|
||||
sdb 100G disk
|
||||
sdc 64G disk
|
||||
sda 32G disk
|
||||
sdb 100G disk
|
||||
sdc 64G disk
|
||||
--- extend the thin pool onto it ---
|
||||
pvcreate ok
|
||||
vgextend ok
|
||||
WARNING: Set activation/thin_pool_autoextend_threshold below 100 to trigger automatic extension of thin pools before they get full.
|
||||
Logical volume pve/data successfully resized.
|
||||
--- reclaim what MY failed pulls left behind (dangling layers only) ---
|
||||
Total reclaimed space: 0B
|
||||
--- after ---
|
||||
LV LSize Data%
|
||||
data <75.81g 15.57
|
||||
root <13.81g
|
||||
swap <3.88g
|
||||
vm-9201-disk-0 32.00g 5.24
|
||||
vm-9201-disk-1 70.00g 14.47
|
||||
TYPE TOTAL ACTIVE SIZE RECLAIMABLE
|
||||
Images 13 5 3.966GB 2.998GB (75%)
|
||||
Containers 5 5 54.78kB 0B (0%)
|
||||
|
||||
## Recovery from MY OWN harness damage — what was changed, and why each change is a fixture change
|
||||
The box was left with a 100 %-full thin pool and nine failed installs. Two fixture changes were made,
|
||||
both to the drill VM, neither to the product:
|
||||
|
||||
1. **A third disk (64 G) was attached to the VM and the LVM thin pool extended onto it.**
|
||||
before: `data` 11.80 g, Data% **100.00**
|
||||
after: `data` 75.81 g, Data% **15.57**
|
||||
The 32 G system disk was simply too small for a twelve-app household once thin-provisioning
|
||||
over-subscribed a 32 G rootfs and a 70 G data volume onto an 11.8 G pool.
|
||||
|
||||
2. **The customer guest's RAM was raised 4096 -> 6144 MB.** The box has 8 GB and the guest had
|
||||
half of it; the controller's memory guard counts COMMITTED memory, so twelve apps against a
|
||||
3712 MB usable budget cannot fit however they are ordered.
|
||||
|
||||
`docker image prune -f` reclaimed **0 B** — the 3 GB `docker system df` calls "reclaimable" are
|
||||
layers still referenced by the 13 pulled images, not dangling ones. Recorded because "75 %
|
||||
reclaimable" reads like free space and is not.
|
||||
|
||||
**These are changes to the drill's own fixture, made in Phase 0 (setup), and they are not part of any
|
||||
round's measurement.** The re-seed that follows installs the apps ONE AT A TIME, waiting for each to
|
||||
reach `deployed: true`, which is what a household does and what the earlier parallel burst was not.
|
||||
=== BEFORE: how is the guest rootfs mounted? ===
|
||||
/dev/mapper/pve-vm--9201--disk--0 on / type ext4 (rw,relatime,stripe=16,emergency_ro)
|
||||
/dev/mapper/pve-vm--9201--disk--1 on /var/lib/docker type ext4 (rw,relatime,stripe=16,emergency_ro)
|
||||
=== restart the guest so ext4 remounts clean (the pool now has room) ===
|
||||
=== AFTER: mount state ===
|
||||
/dev/mapper/pve-vm--9201--disk--0 on / type ext4 (rw,relatime,stripe=16)
|
||||
=== PROOF: can it actually write? (not an assumption) ===
|
||||
WRITE OK
|
||||
=== containers back? ===
|
||||
cloudflared Up 45 seconds
|
||||
felhom-controller Up 44 seconds (healthy)
|
||||
filebrowser Up 43 seconds (healthy)
|
||||
privatebin Up 45 seconds (healthy)
|
||||
traefik Up 45 seconds
|
||||
=== pool ===
|
||||
data <75.81g 15.64
|
||||
vm-9201-disk-0 32.00g 5.24
|
||||
vm-9201-disk-1 70.00g 14.54
|
||||
|
||||
## A filled thin pool wedges the guest READ-ONLY, and adding space does not un-wedge it
|
||||
Worth writing down beyond tonight, because the second half surprised me.
|
||||
|
||||
When the pool hit 100 %, both of the guest's ext4 filesystems remounted themselves with
|
||||
**`emergency_ro`** — visible in the mount flags, not only in dmesg:
|
||||
|
||||
BEFORE:
|
||||
/dev/mapper/pve-vm--9201--disk--0 on / type ext4 (rw,relatime,stripe=16,**emergency_ro**)
|
||||
/dev/mapper/pve-vm--9201--disk--1 on /var/lib/docker type ext4 (rw,relatime,stripe=16,**emergency_ro**)
|
||||
|
||||
The symptom this produced was NOT „no space left": every deploy was refused with
|
||||
„saving app config: writing /opt/docker/stacks/<app>/app.yaml.tmp: **read-only file system**"
|
||||
which reads like a permissions problem and is nothing of the kind. Note the flags still say `rw` —
|
||||
`emergency_ro` sits beside it, so a careless glance at `mount` says the filesystem is writable.
|
||||
|
||||
**Extending the thin pool from 11.8 G to 75.8 G did not clear it.** The space was there (Data% fell
|
||||
to 15.6 %) and every write still failed. It took a guest restart for ext4 to mount clean:
|
||||
|
||||
AFTER: /dev/mapper/pve-vm--9201--disk--0 on / type ext4 (rw,relatime,stripe=16) — no emergency_ro
|
||||
PROOF: `touch /opt/docker/stacks/.rwtest` -> **WRITE OK** (a real write, not an inference from flags)
|
||||
all five containers back in ~45 s; pool 15.64 %
|
||||
|
||||
The proof line matters: „the flags look right" and „the filesystem accepts a write" are different
|
||||
claims, and only the second one is the thing that was broken.
|
||||
|
||||
## CORRECTION — the drive gate was NOT stuck. I was reading a stale snapshot and a silent log.
|
||||
For about four minutes I believed I had found a defect: the data drive was mounted (`df` showed
|
||||
/dev/sdb, 98 G, on both the box and inside the guest) while the controller still recorded
|
||||
`"disconnected": true, "stopped_stacks": ["immich","jellyfin","nextcloud","paperless-ngx"]`, and no
|
||||
`storage_reconnected` event had appeared. I was about to file it.
|
||||
|
||||
**It was self-healing and it healed.** The controller's own DEBUG ring shows:
|
||||
[gate] drive RETURNED /mnt/felhom-drives/hdd_1 — re-attached + restarted gate-stopped apps
|
||||
Event pushed: storage_reconnected (info) — Meghajtó újra csatlakoztatva: Adatlemez
|
||||
and the current state reads:
|
||||
disconnected=None stopped_stacks=None
|
||||
hub event, Sep 16 20:44 info storage_reconnected „Meghajtó újra csatlakoztatva: Adatlemez"
|
||||
|
||||
**Why I nearly got it wrong — two instrument faults at once:**
|
||||
1. `driveGateLoop` runs on a **30-second ticker**; I read the settings file inside that window and
|
||||
treated one sample as a settled state.
|
||||
2. The gate's lines are **DEBUG**, so `docker logs` showed nothing, and I read that silence as
|
||||
„the gate never ran". An absent log line is not evidence — the debug ring had the lines all
|
||||
along (`/api/debug/logs?level=DEBUG`).
|
||||
And a third, smaller one: my first attempt to read the ring parsed the JSON wrongly and reported
|
||||
„total ring entries: 0" for a 29 503-byte response, which looked like confirmation of the silence.
|
||||
|
||||
Recorded in full because the wrong version of this paragraph would have been a filed P-row against a
|
||||
mechanism that works.
|
||||
|
||||
## Tonight's own release, proven through its WHOLE lifecycle on this box (R-543)
|
||||
while escrow_state=pending : the bar was on every page (measured earlier in Phase 0)
|
||||
after the ceremony (escrow_state=**escrowed**), the same four pages:
|
||||
/dashboard escrow-bar-hits=0
|
||||
/launcher escrow-bar-hits=0
|
||||
/backups/apps escrow-bar-hits=0
|
||||
/storage escrow-bar-hits=0
|
||||
The bar appeared while the off-site copy was paused, told the household exactly what to do, and
|
||||
disappeared **for good** when they did it — on a box that installed itself from the published ISO,
|
||||
with no one setting the scene for the test. „védi"/„védené" are both 0 on /backups/apps for now
|
||||
because no class-A app is installed yet; that sentence is checked again once they are.
|
||||
@@ -0,0 +1,147 @@
|
||||
2026-09-16T20:38:31Z login ok (csrf 64)
|
||||
tr: write error: Broken pipe
|
||||
tr: write error: Broken pipe
|
||||
2026-09-16T20:38:31Z bookstack accept=500
|
||||
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/bookstack/app.yaml.tmp: open /opt/docker/stacks/bookstack/app.yaml.tmp: read-only file system
|
||||
tr: write error: Broken pipe
|
||||
2026-09-16T20:38:31Z gokapi accept=500
|
||||
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/gokapi/app.yaml.tmp: open /opt/docker/stacks/gokapi/app.yaml.tmp: read-only file system"}
|
||||
tr: write error: Broken pipe
|
||||
2026-09-16T20:38:31Z homebox accept=500
|
||||
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/homebox/app.yaml.tmp: open /opt/docker/stacks/homebox/app.yaml.tmp: read-only file system"}
|
||||
2026-09-16T20:38:32Z uptime-kuma accept=500
|
||||
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/uptime-kuma/app.yaml.tmp: open /opt/docker/stacks/uptime-kuma/app.yaml.tmp: read-only file sy
|
||||
2026-09-16T20:38:32Z mealie accept=500
|
||||
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/mealie/app.yaml.tmp: open /opt/docker/stacks/mealie/app.yaml.tmp: read-only file system"}
|
||||
tr: write error: Broken pipe
|
||||
2026-09-16T20:38:32Z vaultwarden accept=500
|
||||
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/vaultwarden/app.yaml.tmp: open /opt/docker/stacks/vaultwarden/app.yaml.tmp: read-only file sy
|
||||
tr: write error: Broken pipe
|
||||
tr: write error: Broken pipe
|
||||
2026-09-16T20:38:32Z adventurelog accept=500
|
||||
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/adventurelog/app.yaml.tmp: open /opt/docker/stacks/adventurelog/app.yaml.tmp: read-only file
|
||||
tr: write error: Broken pipe
|
||||
tr: write error: Broken pipe
|
||||
tr: write error: Broken pipe
|
||||
2026-09-16T20:38:32Z nextcloud accept=500
|
||||
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/nextcloud/app.yaml.tmp: open /opt/docker/stacks/nextcloud/app.yaml.tmp: read-only file system
|
||||
tr: write error: Broken pipe
|
||||
tr: write error: Broken pipe
|
||||
tr: write error: Broken pipe
|
||||
2026-09-16T20:38:32Z paperless-ngx accept=500
|
||||
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/paperless-ngx/app.yaml.tmp: open /opt/docker/stacks/paperless-ngx/app.yaml.tmp: read-only fil
|
||||
2026-09-16T20:38:32Z jellyfin accept=500
|
||||
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/jellyfin/app.yaml.tmp: open /opt/docker/stacks/jellyfin/app.yaml.tmp: read-only file system"}
|
||||
tr: write error: Broken pipe
|
||||
2026-09-16T20:38:32Z immich accept=500
|
||||
REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/immich/app.yaml.tmp: open /opt/docker/stacks/immich/app.yaml.tmp: read-only file system"}
|
||||
2026-09-16T20:38:32Z === re-seed finished; final state ===
|
||||
bookstack deployed: false
|
||||
gokapi deployed: false
|
||||
homebox deployed: false
|
||||
immich deployed: false
|
||||
jellyfin deployed: false
|
||||
nextcloud deployed: false
|
||||
paperless-ngx deployed: false
|
||||
privatebin deployed: true
|
||||
uptime-kuma deployed: false
|
||||
vaultwarden deployed: false
|
||||
cloudflared felhom-controller filebrowser privatebin traefik
|
||||
curl: option --data-urlencode: error encountered when reading a file
|
||||
curl: try 'curl --help' or 'curl --manual' for more information
|
||||
grep: /tmp/.l: No such file or directory
|
||||
2026-09-16T20:41:05Z login ok (csrf 0)
|
||||
2026-09-16T20:41:06Z bookstack accept=401
|
||||
REFUSED: {"ok":false,"error":"authentication required"}
|
||||
2026-09-16T20:41:06Z gokapi accept=401
|
||||
REFUSED: {"ok":false,"error":"authentication required"}
|
||||
2026-09-16T20:41:06Z homebox accept=401
|
||||
REFUSED: {"ok":false,"error":"authentication required"}
|
||||
2026-09-16T20:41:06Z uptime-kuma accept=401
|
||||
REFUSED: {"ok":false,"error":"authentication required"}
|
||||
2026-09-16T20:41:06Z mealie accept=401
|
||||
REFUSED: {"ok":false,"error":"authentication required"}
|
||||
2026-09-16T20:41:06Z vaultwarden accept=401
|
||||
REFUSED: {"ok":false,"error":"authentication required"}
|
||||
2026-09-16T20:41:06Z adventurelog accept=401
|
||||
REFUSED: {"ok":false,"error":"authentication required"}
|
||||
2026-09-16T20:41:06Z nextcloud accept=401
|
||||
REFUSED: {"ok":false,"error":"authentication required"}
|
||||
2026-09-16T20:41:06Z paperless-ngx accept=401
|
||||
REFUSED: {"ok":false,"error":"authentication required"}
|
||||
2026-09-16T20:41:06Z jellyfin accept=401
|
||||
REFUSED: {"ok":false,"error":"authentication required"}
|
||||
2026-09-16T20:41:06Z immich accept=401
|
||||
REFUSED: {"ok":false,"error":"authentication required"}
|
||||
2026-09-16T20:41:06Z === re-seed finished; final state ===
|
||||
bookstack deployed: false
|
||||
gokapi deployed: false
|
||||
homebox deployed: false
|
||||
immich deployed: false
|
||||
jellyfin deployed: false
|
||||
nextcloud deployed: false
|
||||
paperless-ngx deployed: false
|
||||
privatebin deployed: true
|
||||
uptime-kuma deployed: false
|
||||
vaultwarden deployed: false
|
||||
cloudflared felhom-controller filebrowser privatebin traefik
|
||||
2026-09-16T20:41:58Z login ok (csrf 64)
|
||||
|
||||
## My THIRD harness error, and the worst-shaped one
|
||||
The guest restart that un-wedged the filesystem also cleared `/tmp` — where the dashboard password
|
||||
file lived. The re-seed script's login therefore failed:
|
||||
|
||||
curl: option --data-urlencode: error encountered when reading a file
|
||||
grep: /tmp/.l: No such file or directory
|
||||
login ok (csrf **0**) <- it said "login ok" with an empty token
|
||||
|
||||
and every one of the eleven deploys came back:
|
||||
|
||||
accept=401 {"ok":false,"error":"authentication required"}
|
||||
|
||||
**Eleven lines that look exactly like the product refusing eleven installs.** They are nothing of the
|
||||
kind: the product was right to refuse an unauthenticated caller, and the missing credential was mine.
|
||||
Had I skimmed this output I would have written up a spectacular false finding — „the box refuses
|
||||
every deploy after a restart" — and it would have been entirely an artefact of my own tooling.
|
||||
|
||||
Three of my errors tonight share one shape: **a script that keeps going after its own precondition
|
||||
failed, and prints a confident line anyway** („all twelve deploys ACCEPTED", „login ok (csrf 0)",
|
||||
and the `head -12` that hid a disk). The fix applied here is the one that should have been there from
|
||||
the first line: the script now ABORTS when the session token is empty, and says the fault is mine
|
||||
rather than reporting deploys as refused.
|
||||
2026-09-16T20:41:58Z bookstack accept=202
|
||||
2026-09-16T20:42:49Z bookstack INSTALLED (containers matching: 2)
|
||||
2026-09-16T20:42:49Z gokapi accept=202
|
||||
2026-09-16T20:42:59Z gokapi INSTALLED (containers matching: 1)
|
||||
2026-09-16T20:42:59Z homebox accept=202
|
||||
2026-09-16T20:43:09Z homebox INSTALLED (containers matching: 1)
|
||||
2026-09-16T20:43:09Z uptime-kuma accept=202
|
||||
2026-09-16T20:44:00Z uptime-kuma INSTALLED (containers matching: 1)
|
||||
2026-09-16T20:44:00Z mealie accept=202
|
||||
2026-09-16T20:44:50Z mealie INSTALLED (containers matching: 1)
|
||||
2026-09-16T20:44:50Z vaultwarden accept=202
|
||||
2026-09-16T20:45:01Z vaultwarden INSTALLED (containers matching: 1)
|
||||
2026-09-16T20:45:01Z adventurelog accept=202
|
||||
2026-09-16T20:46:01Z adventurelog INSTALLED (containers matching: 3)
|
||||
2026-09-16T20:46:01Z nextcloud accept=202
|
||||
2026-09-16T20:46:11Z nextcloud INSTALLED (containers matching: 3)
|
||||
2026-09-16T20:46:12Z paperless-ngx accept=202
|
||||
2026-09-16T20:46:32Z paperless-ngx INSTALLED (containers matching: 0)
|
||||
2026-09-16T20:46:32Z jellyfin accept=202
|
||||
2026-09-16T20:46:42Z jellyfin INSTALLED (containers matching: 1)
|
||||
2026-09-16T20:46:42Z immich accept=202
|
||||
2026-09-16T20:46:52Z immich INSTALLED (containers matching: 4)
|
||||
2026-09-16T20:46:52Z === re-seed finished; final state ===
|
||||
adventurelog deployed: true
|
||||
bookstack deployed: true
|
||||
gokapi deployed: true
|
||||
homebox deployed: true
|
||||
immich deployed: true
|
||||
jellyfin deployed: true
|
||||
mealie deployed: true
|
||||
nextcloud deployed: true
|
||||
paperless-ngx deployed: true
|
||||
privatebin deployed: true
|
||||
uptime-kuma deployed: true
|
||||
vaultwarden deployed: true
|
||||
adventurelog adventurelog-frontend adventurelog-postgres bookstack bookstack-db cloudflared felhom-controller filebrowser gokapi homebox immich-machine-learning immich-postgres immich-redis immich-server jellyfin mealie nextcloud nextcloud-db nextcloud-redis paperless-postgres paperless-redis paperless-webserver privatebin traefik uptime-kuma vaultwarden
|
||||
@@ -0,0 +1,169 @@
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
login ok, csrf len 64
|
||||
tr: write error: Broken pipe
|
||||
tr: write error: Broken pipe
|
||||
tr: write error: Broken pipe
|
||||
tr: write error: Broken pipe
|
||||
tr: write error: Broken pipe
|
||||
2026-09-16T20:30:18Z nextcloud -> 202 {"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
|
||||
tr: write error: Broken pipe
|
||||
2026-09-16T20:30:18Z immich -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá
|
||||
tr: write error: Broken pipe
|
||||
tr: write error: Broken pipe
|
||||
2026-09-16T20:30:18Z bookstack -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá
|
||||
2026-09-16T20:30:19Z privatebin -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá
|
||||
2026-09-16T20:30:19Z gokapi -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá
|
||||
tr: write error: Broken pipe
|
||||
2026-09-16T20:30:19Z vaultwarden -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá
|
||||
tr: write error: Broken pipe
|
||||
tr: write error: Broken pipe
|
||||
2026-09-16T20:30:19Z paperless-ngx -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá
|
||||
2026-09-16T20:30:19Z jellyfin -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá
|
||||
2026-09-16T20:30:19Z mealie -> 400 {"ok":false,"error":"Nincs elég memória az alkalmazás telepítéséhez. Szükséges: 200 MB, Elérhető: 136 MB (öss
|
||||
2026-09-16T20:30:19Z uptime-kuma -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá
|
||||
tr: write error: Broken pipe
|
||||
tr: write error: Broken pipe
|
||||
2026-09-16T20:30:19Z adventurelog -> 400 {"ok":false,"error":"Nincs elég memória az alkalmazás telepítéséhez. Szükséges: 100 MB, Elérhető: 86 MB (össz
|
||||
tr: write error: Broken pipe
|
||||
2026-09-16T20:30:20Z homebox -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá
|
||||
all twelve deploys ACCEPTED (202 = accepted, not installed — R-536: the two are different things)
|
||||
perl: warning: Setting locale failed.
|
||||
perl: warning: Please check that your locale settings:
|
||||
LANGUAGE = (unset),
|
||||
LC_ALL = (unset),
|
||||
LC_CTYPE = "UTF-8",
|
||||
LC_NUMERIC = (unset),
|
||||
LC_COLLATE = (unset),
|
||||
LC_TIME = (unset),
|
||||
LC_MESSAGES = (unset),
|
||||
LC_MONETARY = (unset),
|
||||
LC_ADDRESS = (unset),
|
||||
LC_IDENTIFICATION = (unset),
|
||||
LC_MEASUREMENT = (unset),
|
||||
LC_PAPER = (unset),
|
||||
LC_TELEPHONE = (unset),
|
||||
LC_NAME = (unset),
|
||||
LANG = "en_US.UTF-8"
|
||||
are supported and installed on your system.
|
||||
perl: warning: Falling back to a fallback locale ("en_US.UTF-8").
|
||||
|
||||
## CORRECTION — the line my own script printed is FALSE, and it is corrected here before anything else
|
||||
My seeding script ended with „all twelve deploys ACCEPTED". **Ten were accepted; two were refused.**
|
||||
The sentence was a fixed `echo` at the end of the script, printed without consulting a single result
|
||||
— the exact "a script that announces a conclusion it never checked" failure this project keeps
|
||||
re-learning, and it would have put a false line into tonight's record.
|
||||
|
||||
ACCEPTED (202), 10: nextcloud · immich · bookstack · privatebin · gokapi · vaultwarden ·
|
||||
paperless-ngx · jellyfin · uptime-kuma · homebox
|
||||
REFUSED (400), 2:
|
||||
mealie „Nincs elég memória az alkalmazás telepítéséhez. Szükséges: 200 MB, Elérhető: 136 MB…"
|
||||
adventurelog „Nincs elég memória az alkalmazás telepítéséhez. Szükséges: 100 MB, Elérhető: 86 MB…"
|
||||
|
||||
Note also that nine of the ten acceptances carried a warning of their own:
|
||||
„Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát…"
|
||||
so the box was telling the truth about its memory on the way up, and then refused the last two
|
||||
outright. The refusal is the controller's memory guard doing its job, fail-closed and in plain
|
||||
Hungarian with the numbers in it.
|
||||
|
||||
**Consequence for the drawn schedule, stated now rather than discovered at 23:30:** rounds 1 and 4
|
||||
act on `adventurelog` and `mealie`, and round 6 on `adventurelog` — apps that are NOT installed.
|
||||
The schedule was drawn before the night and is not being re-drawn; what changes is that those rounds
|
||||
must either act on an app that exists, or be recorded as unrunnable. That decision is made and
|
||||
recorded explicitly, not silently.
|
||||
|
||||
Also true and worth keeping: „202 = accepted, not installed" (R-536). How many of the ten actually
|
||||
finished installing is a separate measurement, taken next.
|
||||
|
||||
## The memory facts behind the two refusals — measured
|
||||
the VM (the whole box): 7939 MB total, 5814 MB available
|
||||
the CUSTOMER GUEST (LXC 9201): **memory: 4096, swap: 512, cores: 3**
|
||||
inside the guest at the moment of the refusals: 4096 MB total, ~3685 MB available
|
||||
containers actually running then: 4 (cloudflared, felhom-controller, filebrowser, traefik)
|
||||
|
||||
So the box had ~5.8 GB free while the guest the apps live in was capped at 4 GB — and the guard
|
||||
refused the eleventh and twelfth apps against the guest's cap, counting the memory already COMMITTED
|
||||
by ten in-flight installs rather than the memory currently in use. That is the honest way to count
|
||||
it (otherwise ten simultaneous pulls would all be admitted and then fight), and the message quoted
|
||||
the two numbers it compared.
|
||||
|
||||
**This is the constraint that decides whether tonight's household can be twelve apps at all.** The
|
||||
brief asks for the twelve of BIGNIGHT. The box as installed gives its customer guest 4 GB. Nothing
|
||||
has been changed yet: first the ten in-flight installs are allowed to finish, because the memory
|
||||
picture during a pull is not the memory picture afterwards, and a decision taken on the wrong
|
||||
picture is worse than a late one.
|
||||
|
||||
## DECISION — what happens to the rounds that name apps which are not installed
|
||||
Three rounds name apps the box refused to install: round 1 (`offsite-run adventurelog`), round 4
|
||||
(`offsite-run mealie`), round 6 (`backup-system adventurelog`).
|
||||
|
||||
**The schedule is not re-drawn.** It was fixed from the seed before anything ran, and re-drawing it
|
||||
now — after seeing which apps happened to fit in memory — is exactly the "choose the night after the
|
||||
fact" failure the seed exists to prevent.
|
||||
|
||||
What the rounds actually do, and why this costs less than it looks:
|
||||
|
||||
* **`offsite-run` is a TIER action, not an app action.** The off-site leg is repo-global — one
|
||||
`LastRun` for the whole repository, no per-app run time — so rounds 1 and 4 exercise the tier
|
||||
exactly as drawn. The named app is which app's row I read afterwards; where that app is absent,
|
||||
the round records the tier's own result and says the app was not installed.
|
||||
* **`backup-system` (round 6) is a whole-box action** and does not depend on the named app either.
|
||||
|
||||
So all three rounds run as drawn; what changes is that their "what the customer saw" cell reports the
|
||||
tier or the whole-system page rather than that app's row. Each affected round says so in its own line
|
||||
rather than leaving a reader to assume the app was there.
|
||||
|
||||
**And the memory refusal is itself a finding, not just an inconvenience:** a fresh box built to the
|
||||
documented shape gives its customer guest 4 GB, and the twelve-app household of BIGNIGHT does not fit
|
||||
in it. Nothing was resized to make the drill comfortable.
|
||||
|
||||
## Deploy progress, and a second instrument lesson
|
||||
At 20:32:59Z, twelve minutes after the ten deploys were accepted:
|
||||
app.yaml recorded: 10 (the box registered all ten)
|
||||
containers running: 5 (cloudflared, felhom-controller, filebrowser, traefik, **privatebin**)
|
||||
guest memory available: 3818 MB
|
||||
So exactly one of the ten household apps was actually up; the rest were still pulling images. This is
|
||||
R-536's distinction in the flesh: ten "telepítés elindítva" acceptances, one installed app.
|
||||
|
||||
**The instrument lesson (my second tonight):** I read the deploy watcher with
|
||||
`tail -6 … | grep -vE "locale|…"`, and the locale warnings filled the whole tail, so the watcher's
|
||||
ONE real line was filtered out and the watcher looked dead. I had already declared one watcher dead
|
||||
tonight for a different reason and killed it with a `pkill` that killed itself. The fix both times
|
||||
was the same: **ask the box directly instead of trusting my own reporting layer.** The box answered
|
||||
in one call, and the watcher turned out to have been working the whole time.
|
||||
@@ -0,0 +1,17 @@
|
||||
#!/bin/bash
|
||||
# Run on demo-hp AFTER the PVE install finishes.
|
||||
#
|
||||
# Why this exists: "Automatically reboot after successful installation" is ticked and the boot order
|
||||
# is ide2;scsi0, so the box reboots straight back INTO the installer — and a completed install then
|
||||
# looks exactly like a stuck one. A running guest also keeps the QEMU boot order it started with, so
|
||||
# editing the config mid-run is not enough: the VM must be stopped.
|
||||
#
|
||||
# Completion is judged from the DISK, not from the screen.
|
||||
set -u
|
||||
V=336
|
||||
echo "disk usage before: $(du -sh --block-size=1M /mnt/hdd_1/images/$V/vm-336-disk-1.raw 2>/dev/null | cut -f1) MiB"
|
||||
qm stop $V; sleep 5
|
||||
qm set $V --delete ide2
|
||||
qm set $V --boot order=scsi0 # its OWN call, always
|
||||
qm config $V | grep -E '^(boot|ide2|scsi)'
|
||||
qm start $V && echo "started from disk at $(date -u +%FT%TZ)"
|
||||
@@ -0,0 +1,36 @@
|
||||
#!/bin/bash
|
||||
# round.sh — run ONE chaos round and record the same five things.
|
||||
# Usage: round.sh <n> <action> <app> <accident>
|
||||
set -u
|
||||
N="$1"; X="$2"; Y="$3"; Z="$4"
|
||||
E=/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17
|
||||
OUT="$E/round-${N}.txt"
|
||||
SC=/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/adc14dfe-fc6c-4014-9378-d580d29d3595/scratchpad/chaos
|
||||
say(){ echo "$(date -u +%FT%TZ) $*" | tee -a "$OUT"; }
|
||||
export SSHPASS=$(cat $SC/boxroot.pw)
|
||||
SSHO="-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=8 -o NumberOfPasswordPrompts=1"
|
||||
BOX(){ sshpass -e ssh $SSHO root@192.168.0.115 "$@" 2>/dev/null | grep -v "Warning: Permanently"; }
|
||||
|
||||
say "================ ROUND $N : $X on $Y, while: $Z ================"
|
||||
say "--- BEFORE: is the box steady? (every app up, hub ONLINE, no active alarm) ---"
|
||||
BOX 'pct exec 9201 -- docker ps --format "{{.Names}} {{.Status}}"' | sed 's/^/ /' | tee -a "$OUT"
|
||||
say " events before this round:"
|
||||
bash "$E/events.sh" 3 | tee -a "$OUT"
|
||||
|
||||
say "--- ACTION: $X on $Y ---"
|
||||
# the action itself is driven per-round by the caller's follow-up; this records the start moment
|
||||
say " action start marker"
|
||||
|
||||
say "--- ACCIDENT: $Z (injected 10-60 s after the action starts) ---"
|
||||
if [ "$Z" != "nothing" ]; then
|
||||
bash "$E/inject.sh" "$Z" "$N" | sed 's/^/ /' | tee -a "$OUT"
|
||||
else
|
||||
say " control round — no accident, deliberately"
|
||||
fi
|
||||
|
||||
say "--- AFTER: what the box did by itself ---"
|
||||
BOX 'pct exec 9201 -- docker ps --format "{{.Names}} {{.Status}}"' | sed 's/^/ /' | tee -a "$OUT"
|
||||
BOX 'pvesm status; df -h /mnt/felhom-drives/hdd_1 2>/dev/null | tail -1' | sed 's/^/ /' | tee -a "$OUT"
|
||||
say "--- alarms in this round's window ---"
|
||||
bash "$E/events.sh" 8 | tee -a "$OUT"
|
||||
say "================ END ROUND $N ================"
|
||||
@@ -0,0 +1,37 @@
|
||||
# The per-round record — CHAOS NIGHT
|
||||
|
||||
Every round records the SAME five things, in this order, and nothing is written from memory:
|
||||
|
||||
1. **what the customer saw** — screens quoted verbatim (Hungarian in „ ", searched with ASCII
|
||||
fragments, with a positive and a negative control; accented `grep` dies with a complexity error
|
||||
here, so searching is done in Python)
|
||||
2. **what the box did by itself** — no shell, no help; anything I had to do is an intervention
|
||||
3. **time to steady** — measured from the accident to the moment every app is running, the hub says
|
||||
ONLINE and no alarm is active. **`—` when the box never got there on its own**, never a guess
|
||||
4. **which alarm fired, and was it TRUE**
|
||||
5. **which alarm SHOULD have fired (per 08-alarm-ladder.md) and did not**
|
||||
|
||||
plus **the background loop's failures inside the round's window**, counted from its own log.
|
||||
|
||||
## Two things known BEFORE the night that change how rounds are scored
|
||||
|
||||
- **Rounds 7, 8 and 9 all block the box's network.** An event generated while the hub is unreachable
|
||||
is retried 3 times over ~6 s and then **dropped permanently** (`PushEvent`, no queue). So a missing
|
||||
alarm in those rounds is not evidence the alarm did not fire — it may have been posted into a
|
||||
blocked path. Scored as `LOST-IN-BLOCK`, never as `MISSED`.
|
||||
- **The dedupe is not one window.** 5 minutes applies ONLY to the node/host liveness events; the
|
||||
default operator cooldown is **1 hour**, keyed `customer:type[...]`. A second identical alarm
|
||||
inside an hour is suppressed and written to `notification_log` with status `suppressed` — so
|
||||
"suppressed" and "never fired" are distinguishable, and must be distinguished.
|
||||
|
||||
## Steady-state check (the same command set every round)
|
||||
|
||||
* every app: `docker ps` shows it running AND its front door answers
|
||||
* the hub: the host row reads ONLINE, guests 1/1, agent version present
|
||||
* no active alarm on the dashboard; `notification_log` read for the round's window
|
||||
* the data drive: the storage page reads „Aktív", not „Leválasztva"
|
||||
|
||||
## Evidence discipline
|
||||
|
||||
Evidence is copied off the box **at the end of each round, before the next accident** (R-320) — the
|
||||
one that gets forgotten is the middle one, never the last.
|
||||
@@ -0,0 +1,23 @@
|
||||
seed: 20260917
|
||||
script sha256: 4b98afe65d042df7e7dc417553b33565cfbb4afd451c68456a7cabec0858d2a1
|
||||
|
||||
| # | time | X — the action | Y — the app | Z — the accident |
|
||||
|---|---|---|---|---|
|
||||
| 1 | 23:30 | offsite-run | adventurelog | nothing |
|
||||
| 2 | 23:55 | restore | gokapi | power cut |
|
||||
| 3 | 00:20 | use | bookstack | disk 95% full |
|
||||
| 4 | 00:45 | offsite-run | mealie | tunnel down 10min |
|
||||
| 5 | 01:10 | use | privatebin | docker restarted |
|
||||
| 6 | 01:35 | backup-system | adventurelog | nothing |
|
||||
| 7 | 02:00 | update | nextcloud | internet gone 10min |
|
||||
| 8 | 02:25 | backup-app | nextcloud | internet gone 10min |
|
||||
| 9 | 02:50 | use | uptime-kuma | internet gone 10min |
|
||||
| 10 | 03:15 | restore | uptime-kuma | hard reset |
|
||||
| 11 | 03:40 | use | paperless-ngx | drive pulled 20min |
|
||||
| 12 | 04:05 | use | paperless-ngx | nothing |
|
||||
|
||||
re-draw log (4 entries):
|
||||
r02 X=reinstall re-drawn (nothing has been removed yet)
|
||||
r06 Z=disk 95% full re-drawn (constraint 4: at most once)
|
||||
r07 Z=nothing re-drawn (constraint 6: never two in a row after r2)
|
||||
r08 X=reinstall re-drawn (nothing has been removed yet)
|
||||
|
After Width: | Height: | Size: 25 KiB |
|
After Width: | Height: | Size: 157 KiB |
|
After Width: | Height: | Size: 137 KiB |
|
After Width: | Height: | Size: 131 KiB |
|
After Width: | Height: | Size: 131 KiB |
|
After Width: | Height: | Size: 10 KiB |
|
After Width: | Height: | Size: 134 KiB |
|
After Width: | Height: | Size: 11 KiB |
|
After Width: | Height: | Size: 136 KiB |
|
After Width: | Height: | Size: 11 KiB |
|
After Width: | Height: | Size: 137 KiB |
|
After Width: | Height: | Size: 150 KiB |
|
After Width: | Height: | Size: 133 KiB |
|
After Width: | Height: | Size: 43 KiB |
|
After Width: | Height: | Size: 9.7 KiB |
|
After Width: | Height: | Size: 11 KiB |
|
After Width: | Height: | Size: 22 KiB |
@@ -0,0 +1,47 @@
|
||||
#!/bin/bash
|
||||
# seed_apps.sh — CHAOS NIGHT Phase 0.3: move the household in.
|
||||
# Twelve apps, deployed through the controller's own API (the endpoint the UI invokes), each with
|
||||
# exactly the fields its template declares. Secrets are GENERATED here and written to a 0600 file;
|
||||
# none is ever echoed. Run INSIDE the guest (the controller answers on the container address).
|
||||
#
|
||||
# Usage: seed_apps.sh <controller-container-ip> <host-header> <pw-file> <hdd-path> <secrets-out>
|
||||
set -u
|
||||
B="http://${1:?ctrl ip}:8080"; H="Host: ${2:?host header}"; PWF="${3:?pw file}"
|
||||
HDD="${4:?hdd path}"; SEC="${5:?secrets out}"
|
||||
: > "$SEC"; chmod 600 "$SEC"
|
||||
gen(){ tr -dc 'a-zA-Z0-9' </dev/urandom | head -c "${1:-24}"; }
|
||||
|
||||
curl -s -o /dev/null -D /tmp/.l -H "$H" -X POST --data-urlencode "password@$PWF" $B/login
|
||||
SESS=$(grep -i '^set-cookie: felhom_session=' /tmp/.l | sed 's/.*felhom_session=\([^;]*\).*/\1/')
|
||||
[ -n "$SESS" ] || { echo "LOGIN FAILED"; exit 1; }
|
||||
C="Cookie: felhom_session=$SESS"
|
||||
TOK=$(curl -s -H "$H" -H "$C" $B/backups/remote | grep -o 'name="csrf-token" content="[^"]*"' | sed 's/.*content="\([^"]*\)".*/\1/')
|
||||
echo "login ok, csrf len ${#TOK}"
|
||||
|
||||
dep(){ # dep <stack> <json-values>
|
||||
local app="$1" vals="$2"
|
||||
local code
|
||||
code=$(curl -s -o /tmp/.d -w '%{http_code}' -H "$H" -H "$C" -H "X-CSRF-Token: $TOK" \
|
||||
-H 'Content-Type: application/json' -X POST -d "{\"values\":$vals}" \
|
||||
"$B/api/stacks/$app/deploy")
|
||||
printf '%s %-14s -> %s %s\n' "$(date -u +%FT%TZ)" "$app" "$code" "$(head -c 120 /tmp/.d)"
|
||||
}
|
||||
|
||||
D='"DOMAIN":"enkicsifelhom.hu"'
|
||||
NC_ADMIN=$(gen 20); PL_ADMIN=$(gen 20); GK=$(gen 20)
|
||||
{ echo "nextcloud_admin=$NC_ADMIN"; echo "paperless_admin=$PL_ADMIN"; echo "gokapi_admin=$GK"; } >> "$SEC"
|
||||
|
||||
dep nextcloud "{$D,\"SUBDOMAIN\":\"cloud\",\"DB_PASSWORD\":\"$(gen)\",\"MYSQL_ROOT_PASSWORD\":\"$(gen)\",\"NEXTCLOUD_ADMIN_USER\":\"admin\",\"NEXTCLOUD_ADMIN_PASSWORD\":\"$NC_ADMIN\",\"HDD_PATH\":\"$HDD\"}"
|
||||
dep immich "{$D,\"SUBDOMAIN\":\"photos\",\"DB_PASSWORD\":\"$(gen)\",\"HDD_PATH\":\"$HDD\"}"
|
||||
dep bookstack "{$D,\"SUBDOMAIN\":\"wiki\",\"APP_KEY\":\"base64:$(gen 32)\",\"DB_PASSWORD\":\"$(gen)\"}"
|
||||
dep privatebin "{$D,\"SUBDOMAIN\":\"paste\"}"
|
||||
dep gokapi "{$D,\"SUBDOMAIN\":\"share\",\"GOKAPI_PASSWORD\":\"$GK\"}"
|
||||
dep vaultwarden "{$D,\"SUBDOMAIN\":\"vault\",\"ADMIN_TOKEN\":\"$(gen 32)\",\"SIGNUPS_ALLOWED\":\"false\"}"
|
||||
dep paperless-ngx "{$D,\"SUBDOMAIN\":\"paperless\",\"DB_PASSWORD\":\"$(gen)\",\"PAPERLESS_SECRET_KEY\":\"$(gen 32)\",\"PAPERLESS_ADMIN_USER\":\"admin\",\"PAPERLESS_ADMIN_PASSWORD\":\"$PL_ADMIN\",\"HDD_PATH\":\"$HDD\",\"PAPERLESS_OCR_LANGUAGE\":\"hun+eng\"}"
|
||||
dep jellyfin "{$D,\"SUBDOMAIN\":\"media\",\"HDD_PATH\":\"$HDD\"}"
|
||||
dep mealie "{$D,\"SUBDOMAIN\":\"recipes\"}"
|
||||
dep uptime-kuma "{$D,\"SUBDOMAIN\":\"status\"}"
|
||||
dep adventurelog "{$D,\"SUBDOMAIN\":\"travel\",\"SECRET_KEY\":\"$(gen 32)\",\"DB_PASSWORD\":\"$(gen)\"}"
|
||||
dep homebox "{$D,\"SUBDOMAIN\":\"inventory\",\"HBOX_AUTH_API_KEY_PEPPER\":\"$(gen 32)\"}"
|
||||
rm -f /tmp/.l /tmp/.d
|
||||
echo "all twelve deploys ACCEPTED (202 = accepted, not installed — R-536: the two are different things)"
|
||||
@@ -0,0 +1,45 @@
|
||||
#!/bin/bash
|
||||
# steady.sh — is the box back on its own? Run between every round.
|
||||
#
|
||||
# "Steady" is not a word here, it is a measurement: every app running, the hub says ONLINE, the data
|
||||
# drive reads Aktiv, and no alarm is active. If the box did not get there BY ITSELF, the round's
|
||||
# steady-state cell is `—`, never a guess and never a number I helped it reach.
|
||||
#
|
||||
# Usage: steady.sh <round-label> (reads the box address from $BOXIP, guest id from $GUEST)
|
||||
set -u
|
||||
R="${1:?round label}"
|
||||
VM=336
|
||||
E=/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17
|
||||
OUT="$E/round-${R}-steady.txt"
|
||||
H(){ ssh hp "$@" 2>/dev/null | grep -vE "locale|LC_|LANG|perl:|supported and installed|are supported"; }
|
||||
say(){ echo "$(date -u +%FT%TZ) $*" | tee -a "$OUT"; }
|
||||
|
||||
say "=== steady check, round $R ==="
|
||||
say "--- the box itself (is the VM even up?) ---"
|
||||
H "qm status $VM" | tee -a "$OUT"
|
||||
|
||||
say "--- the customer guest and its apps (asked of the box, not of me) ---"
|
||||
H "qm guest exec $VM -- bash -c 'pct list; echo ---; pct exec 9000 -- docker ps --format \"{{.Names}} {{.Status}}\"'" 2>/dev/null | tee -a "$OUT"
|
||||
|
||||
say "--- the hub's view: host row, guests, agent, last report ---"
|
||||
cd /mnt/5_hdd/felhom.eu/git/felhom.eu
|
||||
python3 scripts/read_credential.py HUB_PW /tmp/.shp >/dev/null 2>&1
|
||||
HUB_PW=$(cat /tmp/.shp); IP=$(sudo kubectl -n felhom-system get svc hub -o jsonpath='{.spec.clusterIP}')
|
||||
curl -s -u ":$HUB_PW" "http://$IP:8080/hosts" | python3 -c "
|
||||
import sys,re,html
|
||||
t=sys.stdin.read()
|
||||
txt=html.unescape(re.sub(r'\s+',' ',re.sub(r'<[^>]+>',' ',re.sub(r'<script.*?</script>','',t,flags=re.S))))
|
||||
i=txt.find('chaosnight')
|
||||
print(' ', txt[max(0,i-40):i+220] if i>=0 else 'chaosnight NOT in the hosts table')
|
||||
" | tee -a "$OUT"
|
||||
|
||||
say "--- alarms/events in the last 30 minutes (fired AND suppressed are different things) ---"
|
||||
curl -s -u ":$HUB_PW" "http://$IP:8080/" | python3 -c "
|
||||
import sys,re,html
|
||||
t=sys.stdin.read()
|
||||
txt=html.unescape(re.sub(r'\s+',' ',re.sub(r'<[^>]+>',' ',re.sub(r'<script.*?</script>','',t,flags=re.S))))
|
||||
m=re.search(r'(Recent events|Events|Legutobbi).{0,900}', txt)
|
||||
print(' ', (m.group(0)[:900] if m else txt[:400]))
|
||||
" | tee -a "$OUT"
|
||||
rm -f /tmp/.shp
|
||||
say "=== end steady check, round $R ==="
|
||||
@@ -728,6 +728,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-543** | **[P1-HIGH] Off-site ON by default is not off-site WORKING: on a fresh box tier 3 sits at „Kulcsletétre vár" until the household does the escrow ceremony, and nothing asks them to — while the tier-1 row now tells them their files are protected by that very copy.** MEASURED 2026-09-16 on the fresh box (controller 0.244.0, hub 0.116.0, off-site provisioned automatically by the new default): the app-backup page reads „3. mentés — Kulcsletétre vár · A távoli mentés a titkosítási kulcs letétbe helyezéséig szünetel", the remote page reads „Helyreállítási kód szükséges", and `POST /backup/offbox/run` returns 302 while producing no snapshot (the controller log shows only `offsite-credential-retry`, no restic activity). **Why it matters more than before today:** hub v0.116.0 makes off-site the default *because* a one-drive box otherwise keeps the household's files in no tier at all (R-537/R-538), and controller v0.244.0 now prints „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi" under the tier-1 row. On day one both are true-in-intent and false-in-fact: the copy is paused. **Fix shape (one of):** prompt the escrow ceremony as part of first-run when off-site is enabled and un-escrowed; and/or make the tier-1 sentence state the tier's actual state („…védené — a távoli mentés a helyreállítási kód létrehozásáig szünetel"). The ceremony itself works and is customer-facing („Helyreállítási kód létrehozása"); what is missing is that anyone is told to do it. **CLOSED 2026-09-16 — controller v0.245.0, both halves proven live.** The pause is untouched: it is the zero-knowledge escrow design, and this row was never about the mechanism. (a) **The household is asked:** while the off-site tier is configured and its escrow is not complete, every authenticated page carries „A távoli mentés szünetel, amíg nem hozod létre a helyreállítási kódot." linking `/backup/escrow`. It is the R-241 bar, second instance — same session-cookie dismissal, back at the next visit, gone for good when escrowed; **no second banner system**. It hangs off `executeTemplate`, the single render choke point, so it cannot reach only the pages someone remembered. (b) **The tier-1 sentence renders by state:** `driveFilesNoteFor` takes `tier3State`'s own vocabulary — `active` → „védi", `escrow_pending` → „védené … a helyreállítási kód létrehozásáig szünetel" + the route, no off-site and no second drive → „nincs másolat" + both ways out. **Measured live on 0.245.0:** on a paused box (9202, off-site configured through the product's own endpoint, `escrow_state=pending`) the bar renders on /dashboard, /launcher, /backups/apps and /settings; a manual `POST /backup/offbox/run` is refused by the fork-4 gate („A távoli mentés a kulcs letétbe helyezésére vár.") with `last_run=None, snapshot_count=None`; and a throwaway class-A app's row reads „…védené … szünetel" with „védi"=0. On an escrowed box (9201) the bar is absent on all three pages and the row reads „védi". Both red-proofed (the bar test fails on BOTH pages with the one hook line removed; the sentence test quotes the exact v0.244.0 promise when the state is ignored). The first-hour guide now asks for the code right after the dashboard password and before the first app. Evidence: `audits/evidence-recovery-code-2026-09-16/`. | **CLOSED 2026-09-16 — shipped in controller v0.245.0 and proven live** |
|
||||
| **R-544** | **[P3-LOW] The host-delete log line says „escrow deleted: true" while the documented (and actual) effect is DEMOTION to retained custody.** MEASURED 2026-09-16 during the teardown of the fresh box: an unacknowledged delete was correctly refused 409 („has key escrow (acknowledgement missing)") and the refusal text promises the acknowledgement „moves it to retained custody"; the acknowledged delete then logged `host deleted: tester-1-33b6a9 (escrow deleted: true)`. The hub's own customer page states the truth — „host deletion only demotes custody, never destroys it… recovery-key custody is demoted to retained custody, not destroyed", with the customer delete named as „the one true purge point". **Nothing is broken; the log is.** An operator reading that line during an incident would believe a household's last key had just been destroyed, and the R-304 retention exists precisely so it is not. **Fix shape:** log what happened — `escrow custody demoted to retained (host delete)` — and keep the boolean's name out of operator-facing text. | **READY — rank P3-LOW; owner: CC (hub)** |
|
||||
| **R-545** | **[P3-LOW] There is no product action that un-configures an off-site target — only one that re-starts an ORPHANED repository.** FOUND 2026-09-16 while exercising R-543's paused state on the scratch guest: `POST /backup/offbox/config` configures a target and can disable it (`enabled` unchecked), but nothing removes it. `POST /backup/offbox/reset` refuses unless `OffboxOrphaned()` is true („Az offsite tároló nincs elárvult állapotban."), and it means „start a new remote backup, set the old history aside" — not „forget this destination". So a household that sets up the wrong NAS, or a box being handed to someone else, keeps the host, user, path, ssh key and minted repo password on disk with no route to clear them; a disabled target still holds its secrets in `data/offbox/`. **Why P3 and not higher:** a disabled target runs nothing and the secrets are 0600 on the box's own disk, so nothing leaks and no copy is lost. **Fix shape:** a „Távoli cél törlése" action beside the config form that clears the target and shreds `data/offbox/`, REFUSING while the hub holds a sealed package for this box (the R-241 rule — dropping the key would orphan the history that package protects). Teardown for this session's proof had to clear it out-of-band for exactly this reason, which is the measurement. | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-546** | **[P2-MEDIUM] The first-hour guide sends the household to create their recovery code at a moment when the box cannot yet do it — and the new reminder bar urges them there on every page.** MEASURED 2026-09-16/17 on a fresh box (`tester-1-022354`, guest 9201, controller 0.245.0, agent 0.131.0, installed from the published ISO 1.28.0). `VOLUNTEER-first-hour.md` §6 — added hours earlier in controller v0.245.0 — places „A helyreállítási kód" immediately after the dashboard password and **before the first app**, because until it is done the off-site copy does not run. At exactly that point the ceremony FAILS: `POST /api/escrow/start` → 200, then `GET /api/escrow/status` → `detail: "exit 2: … selftest=escrow-create requires -storage <pbs-storage-id> (or escrow.pbs_storage…"`, and `POST /api/escrow/claim` → **409** „A folyamat jelenlegi állapotában a kód nem kérhető le." **Cause, measured on both sides:** the hub had auto-provisioned the DR descriptor at 20:19 (no press — see R-534/R-511's acknowledged-delete path) and its Backup & DR panel itself read „descriptor provisioned … **waiting** · ceremony possible once the descriptor is applied on the box"; the box had no PBS storage (`pvesm status` = local + local-lvm only) and `/etc/felhom-agent/agent.json` had **no `escrow` section at all**. **It is a TIMING gap and it self-heals:** a watcher left the box alone and polled — `pbs_storage` and `escrow.pbs_storage_id` both became `felhom-pbs` at **20:35:16Z, ~17 minutes after the bind**; the retried ceremony then passed every preflight item and the claim returned 200 (83-character code, entropy 129.2 bits), and `escrow_state` flipped to `escrowed`. **Why it still matters:** for those ~17 minutes the R-543 reminder bar (also v0.245.0) is on *every* page telling the household to do the one thing that refuses, and nothing on the page says „wait a few minutes" — the volunteer meets a stderr fragment about a `-storage` flag. **Fix shape (one of):** have the escrow page/bar consult `preflight` and say „a doboz még készül — pár perc múlva próbáld újra" while `pbs_storage_id` is unset; or move the guide's step to after the first app; or make the bar appear only once preflight is green. **No product code was changed tonight** (validation run). Evidence: `audits/evidence-chaos-night-2026-09-17/phase0-escrow-failure.txt` and `phase0-escrow-retry.txt`. | **READY — rank P2-MEDIUM; owner: CC (controller copy + guide timing)** |
|
||||
| **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured, and a class-A app carries one sentence saying where its files ARE protected. Proven on demo-hp through the page the customer opens: Paperless-ngx reads „1. mentés … DB + Konfig" with „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi …", while its „2. mentés" row still reads „DB + Konfig + Adatok". Red-proof: restoring the old app-shaped label fails `TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold`. **RE-PROVEN 2026-09-16 on a FRESH box** (installed from the built ISO 1.28.0, controller 0.244.0, off-site on by default): the Nextcloud row read „1. mentés … DB + Konfig" with the new sentence, „2. mentés … Nincs 2. (off-drive) másolat", „3. mentés Sikeres restic → …your-storagebox.de"; „DB + Konfig + Adatok" appeared ZERO times while the local unit held no file leg. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** |
|
||||
| **R-538** | **[P1-HIGH] A tier-1 app restore reports plain success and leaves Nextcloud listing files whose bytes were never in the backup — and it destroys the app's own trash, the customer's last copy.** MEASURED 2026-09-16 on the drill box, F10 („a child deletes the photo folder"): the five photos were deleted through Nextcloud (DELETE 204, PROPFIND 404), then restored through the page exactly as a customer would (`POST /backup/restore` `stack_name=nextcloud` `snapshot_id=helyi` → 302, finished in **35 s**, „A(z) nextcloud: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult."). Afterwards the folder is back and **lists all five photos**, and **none of them opens**: `GET nyaralas-1..5` = 404 / 503×4 with `Sabre\DAV\Exception\NotFound`, while the positive controls at the same moment pass (`status.php` 200, WebDAV PUT 201, GET 200). Cause: the replayed MariaDB dump (11:01:45Z) knows the photos, the bytes live on `mp8` and were never captured (R-537). **Worse:** the bytes were still on the drive in Nextcloud's own trash (`appdata/nextcloud/admin/files_trashbin/files/Fotok.d1789556707/nyaralas-1..5.jpg`, all five present) and the restored database no longer references them — the trash listing comes back **empty**, so „restore from trash", the one route that would have worked, is gone. The customer is left with five unopenable photos, a success message, and no warning. **Fix shape:** before replaying a database whose app has an uncaptured file leg, refuse or warn („ennek az alkalmazásnak a fájljai nincsenek ebben a mentésben — a visszaállítás után a fájlok hiányozni fognak"); and never present a DB-only restore of a class-A app as a complete one. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** A unit restore refuses before anything is touched when the unit cannot return the app's drive-side files, and names the route that can. Fired live on demo-hp: `POST /backup/restore` for paperless-ngx → 302 with „Ez a mentés nem tartalmazza az alkalmazás fájljait, ezért nem állítjuk vissza az adatbázist föléjük — a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza …", and the app read `running` before AND after, so nothing was stopped and no trash was made unreachable. The database-and-settings-only path exists as a separately worded second step. Red-proof: disabling the guard fails `TestUnitRestore_RefusesWhenTheUnitCannotHoldTheFiles`. **RE-PROVEN 2026-09-16 on a FRESH box, and this time the refusal had somewhere to point:** after five photos were deleted, `POST /backup/restore` was refused with „…a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza: … „Teljes visszaállítás (fájlok + adatbázis)"", the app read `running` before AND after, and the wastebasket was untouched. The off-site route then returned all five photos — 200 with the exact uploaded sizes and sha256 IDENTICAL to the originals, 5/5, with a negative control. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseE-photos.txt`. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** |
|
||||
| **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** |
|
||||
|
||||