## CHAOS NIGHT — Phase 0 notes (2026-09-16 evening, CEST)

### Baselines, re-verified live against Gitea at 21:49 CEST (not copied from the brief)
felhom-controller 714d5bce0920 v0.245.0 (MinAgent 0.131.0)   — clean, in sync
felhom-agent      e98b857684f4 v0.131.0                      — clean, in sync
felhom.eu         d124c77e176d hub v0.116.0, ISO 1.28.0 live — clean, in sync
app-catalog       94bc5febaca2                                — clean, in sync
Register: highest R-545, 212 open. Golden waiver valid to 2026-09-27.

### A claim in the brief, CHECKED rather than inherited
The brief says the customer `tester-1` has "no host". CONFIRMED: /hosts lists exactly three hosts —
demo-felhom-8363b5, demo-hp-bb76ea, drill-r50-0a4f9a. There is no tester-1 host record. The
customers list showing „tester-1 … DOWN … 0.244.0" is the customer's LAST KNOWN state, not a live
box; reading that row as a host record would have been the mistake.

### 0.1 — golden 0.245.0, baked and published
Launched 19:52:47Z as a transient unit in the drill VM; finished 19:58:15Z.
  markers: overlay2=1  mountpoints=2  upload=1  FATAL=0  publish-skipped=0
  GOLDEN_VERSION=0.245.0
  GOLDEN_SHA256=7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626
  REGISTRY CHECK (the outcome, not the attempt): golden 0.245.0 -> http=200
  then: guest 9100 destroyed, token shredded, qemu exited, disk reverted to virgin.

  All three of 2026-09-16's bake failures were guarded against and none recurred: `scp -P` (not the
  ssh `-p`), `chmod 0700` on the script, and `GITEA_USER=admin` (not the first credentials line).
  The token-leak check ran with a needle PROVEN non-empty first, because `grep -F ""` matches every
  line and an instrument that reports a hit on an empty needle is not a measurement.

  HONEST NOTE ON THE EXIT CODE: the wrapper script exited **144** while every measured outcome was
  good. That is why teardown was gated on the REGISTRY answering 200 and not on an exit code — this
  repo's own "exit codes that lie" class. The artifact is published and verified independently.

### 0.1b — vouched in the hub
POST /configuration/artifacts -> 303, then READ BACK from the page (the outcome, not the POST code):
  golden currently vouched: 0.245.0
  golden sha now:           7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626
  agent 0.131.0 and the wrapper sha left exactly as they were.

DECISION, and the reason, because silence reads as agreement: the GLOBAL controller floor was left at
0.244.0 and NOT raised to 0.245.0. The new box installs golden 0.245.0, which already carries
controller 0.245.0, so the floor is not needed to deliver anything tonight; raising it would have
pushed a controller update onto demo-felhom, a box that is not part of this drill, at 22:00 at night.
The brief asked for a bake and a vouch, not a floor raise.

### 0.2 — the box
demo-hp (Tier 0). VM 336 „tester1-chaos-night": 8 GiB, 4 cores, cpu host, virtio-scsi-single,
scsi0 = 32 G system disk, scsi1 = 100 G data disk, both on `nvme-scratch` (dir storage, path
/mnt/hdd_1, is_mountpoint yes — the NVMe at its ROOT, per target-selection.md). CD-ROM is the
PUBLISHED felhom-installer-1.28.0-pve9.2-1.iso. **The boot order was set in its own `qm set`** —
combining it with the disk call silently yields `order=net0;ide2` and the VM netboots.
  boot: order=ide2;scsi0   (verified by reading `qm config 336` back)
The VM took DHCP 192.168.0.115 and booted the installer's graphical entry.

### Console driving — measured, not assumed
  * Enter on the EULA page: advances.
  * Enter on the Location page: lands IN the Country text field and does NOT press Next (the
    documented GTK trap).
  * **Alt+N works as the „Next" mnemonic** and is what actually drives this installer. Recorded
    because the previous drill switched to the text-mode entry to avoid exactly this problem.
  * The VM has a QEMU HID Tablet (absolute), so `mouse_move <x> <y>` takes absolute coordinates and
    a real click is available as a fallback. `info mice` says so — checked, not assumed.
  * The installer REFUSES the prefilled `mail@example.invalid` with „Please enter a valid Email
    address" and simply does not advance. The dialog is above the fold, which is why the cropped
    strip looked like „nothing happened" — the full screen showed the reason.
  * Keyboard layout is Hungarian (QWERTZ): `y` and `z` are swapped and `@` is AltGr+V. The drill root
    password was generated from a-x plus digits so it is identical under either layout, and it is
    stored out-of-band (0600, scratchpad) — never in a committed file.

### Console driving on a Hungarian-layout installer — MEASURED, and it cost four round trips
These are facts about driving a PVE 9.2 graphical installer headlessly through `qm monitor`, and
every one of them was measured on this box tonight rather than recalled:

  * `sendkey alt-n` is the „Next" mnemonic and is what actually advances the installer. Plain `ret`
    advances ONLY the EULA page; on every later page it lands inside a text entry and does nothing.
  * **`sendkey altgr-v` does NOTHING.** On a Hungarian layout `@` is AltGr+V, and the QEMU key name
    that works is **`alt_r`** — `sendkey alt_r-v` types the `@`. The failure is silent: the character
    is simply absent, so `admin@felhom.eu` became `adminfelhom.eu` and the installer then refused the
    page with „Please enter a valid Email address" — a refusal that looks exactly like „nothing
    happened" if you only crop the bottom strip of the screen.
  * `mouse_move <x> <y>` did not move the pointer even though `info mice` reports a QEMU HID Tablet
    (absolute) as the active device. Clicking was abandoned; keyboard navigation is the reliable path.
  * **Focus is found by MEASURING it, not by counting tabs.** A 20-line script samples the blue border
    of each entry box in the screendump and prints which one is focused; Tab is then pressed until
    the wanted field reports focus. Counting keystrokes blind is how a password ends up in the wrong
    field. (`blueness = mean(B-R)` over the box's top border: focused +58, unfocused 0.)
  * The installer REFUSES the prefilled `mail@example.invalid`.

### Fence check at this point
demo-hp: /mnt/hdd_1 has 876 G free; VM 336's two raw disks are sparse (132 G apparent, 12 K actual).
`/` on demo-hp is at 77 % and was deliberately NOT used for VM disks. Guests 9201 and 9202 untouched
and running (they are the standing demo boxes, and 9202 is the scratch guest the morning restore
will use). No `local-lvm`, no prune, `drill-r50` not touched.

### A second brief claim, CHECKED rather than inherited: „the automatic mail is waiting in the mailbox"
CONFIRMED. The mailbox holds three „[Felhom] Kösd össze a Felhom dobozodat" messages to
`tester1@felhom.eu` from `monitoring@felhom.eu`, the newest at **2026-09-16T18:17:46Z** — the same
second as yesterday's host delete, which is R-509's automatic trigger firing. So the box installed
tonight should be bindable with **no operator press**, and the pre-declared intervention O1 (the
„Send self-bind link" button) should not be needed. Whether it IS needed is a measurement of this
night, not an assumption: if the waiting link turns out to be superseded or refused, that is a
finding and the press becomes O1.

The other mail in that mailbox worth noting, because it is the same customer's history and could
confuse a reader of this evidence: two „Új beállító kód — újratelepült a szervered" messages
(10:01:57Z and 16:00:58Z) and one „Beállító kód a jelszavad visszaállításához" (11:12:06Z), all from
2026-09-16 — those belong to yesterday's drills, not to tonight's box.

NOT recorded here, deliberately: the bind link itself. It carries a one-time token, and a one-time
secret does not go into a committed file (nor was the mail body fetched into the session transcript
for the same reason — the link will be taken from the hub at bind time instead).

### The one catalog bump — what it is, and why the ORDER matters
Round 7 drew `update nextcloud`, so the bump has to be on the **nextcloud** template, not on whatever
small app I would have picked. The brief asked for "one drill bump of one SMALL app"; the draw
overrides the choice of app, so the compromise is to bump the template's small sidecar pin rather
than the big application image:

  templates/nextcloud/docker-compose.yml:103   redis:7-alpine  ->  redis:7.4-alpine

`redis:7.4-alpine` was CHECKED to exist on Docker Hub before being written down — inventing a tag
would have made round 7 fail for the wrong reason, and a round that fails for the wrong reason
measures nothing. `catalog_since` moves with it, per this repo's own rule that any commit changing an
`image:` line must.

**The bump is NOT pushed yet, and the order is the point:** nextcloud must be DEPLOYED from the
current catalog first, or the box installs 7.4-alpine immediately and round 7 has no update to
apply. Sequence: seed the twelve apps → then push the bump → the box's git-sync picks it up within
15 min (or the „Sablonok frissítése" button) → round 7 at 02:00 has a real update waiting.

Fence note: the catalog is SHARED with the demo boxes. A bump offers an update; it never applies one
— the guarded update needs a press — so the demo boxes' standing apps are not disturbed by this.

### Deploy fields, read from the templates rather than guessed
All twelve templates live under `templates/<app>/`, not at the repo root (my first lookup used the
wrong layout and returned "NO .felhom.yml" twelve times — recorded because the wrong answer was
uniform and therefore looked authoritative). Only SUBDOMAIN and HDD_PATH are marked `required`;
every secret field is optional in metadata but the server refuses a deploy without them (the
metadata-vs-server contradiction measured 2026-09-16), so the seeding script supplies generated
values for all of them and writes them to a 0600 file that is never echoed.
Apps needing the data drive: nextcloud, immich, paperless-ngx, jellyfin.

### 0.2b — the install itself
Install pressed 20:07:11Z. **Completion was judged from the DISK, not from the screen** — a PVE
install with "Automatically reboot" ticked re-enters the installer, so a finished install and a stuck
one look identical on screen. The system disk's blocks-used grew 3233 → 6942 MiB and then held
steady across three consecutive 30-second checks:
  20:09:36Z 6942 MiB (stable 0) · 20:10:07Z 6942 (1) · 20:10:37Z 6942 (2) · 20:11:08Z 6942 (3)
  -> INSTALL COMPLETE 20:11:08Z

Then the documented post-install fix, because a RUNNING guest keeps the boot order QEMU started
with — editing the config mid-run is not enough:
  qm stop 336 · qm set 336 --delete ide2 · qm set 336 --boot order=scsi0 (its own call) · qm start
  read back: `boot: order=scsi0`, no ide2, both disks intact
  started from disk 20:11:19Z

The summary page was checked before Install was pressed, and the line that mattered was
**„Disk(s): /dev/sda"** — the 32 G system disk alone. The 100 G data disk was NOT offered to the
installer and was not touched. That is the check that makes pressing Install safe on a box with a
second disk, and it is the one a two-disk filter has silently got wrong elsewhere in this project.

### How this box is driven, and how the alarms are read — method facts, each one measured
  * **`qm guest exec` is UNUSABLE on this box.** The VM was created with `--agent 1`, but a plain PVE
    install does not run `qemu-guest-agent`, and `qm agent 336 ping` answers nothing. My first
    first-boot watcher was written against `qm guest exec` and therefore sat silent while reporting
    nothing — it looked like a box that would not boot, and it was an instrument pointed at nothing.
  * **The access path is SSH to the box** (`root@192.168.0.115`, the drill password via `sshpass -e`
    from a 0600 file, never on a command line). Confirmed with `LOGIN_OK`, hostname `chaosnight`,
    `pve-manager/9.2.2`.
  * **The alarm instrument is the hub's own `events-table`** on `/customers/tester-1` (note: the
    customer page is `/customers/<id>`, NOT `/configs/<id>` — the latter redirects). The tabs are
    client-side, so `?tab=events` returns the same document; the events must be pulled out of the
    table by its container id. The hub pod has **no `sqlite3`** and `hub.db` is 298 MB with a live
    `-wal`, so copying the DB is both heavy and stale-prone; the page is the correct instrument.
  * **Events baseline, captured before round 1:** the newest pre-drill event is
    `Sep 16 18:47 error node_down` (yesterday's box). Anything newer belongs to tonight. Without this
    marker, yesterday's `node_stale`/`node_down`/`selfbind_link_sent` rows would be scored as
    tonight's alarms.

### My own mistake, recorded because it cost two watchers
`pkill -f "first-boot watcher"` **matched its own command line** and killed the very background job
it was meant to clear, twice, each exiting 144. Same class as the documented `pgrep -f
qemu-system-x86_64` self-match. The replacement watcher does no pkill at all.

### Fence note on a secret
While following redirects to find the customer page, `curl -w '%{url_effective}'` printed the hub
password back to me inside the resolved URL. It is not in any file written here, and that format
option is not used again. Recorded rather than quietly dropped, because the next person will hit the
same flag.

### 0.2c — the bind, as a volunteer does it: ZERO operator presses
The box registered itself as an unclaimed appliance and then sat polling every 30 s — correctly, and
visibly: „not bound yet — polling every 30s until the operator or a customer self-bind lands". No
agent and no guest exist until the bind happens, so a reader who expected the box to finish
installing by itself would have mis-read a waiting box as a stuck one.

The volunteer's own path was taken, end to end:
  * the link came from the **waiting mail** (minted 2026-09-16T18:17:46Z by yesterday's host delete,
    R-509's automatic trigger) — not from the operator's „Send self-bind link" button;
  * the **pairing code** was read off the box's own console: `4SY-4TX`;
  * the **„Tulajdonosi jelmondat"** (5 words) came from the hub's customer record, which is where the
    operator hands it from — it is never e-mailed, by design.

  POST /bind/<token> -> 200 at 2026-09-16T20:18:15Z
  „Sikeres összekötés. A doboz kb. egy percen belül folytatja a telepítést. Ezt az oldalt
   bezárhatod — a beállítás a háttérben befejeződik, és a vezérlőpultod hamarosan elérhető lesz."

**O1 (pre-declared) was NOT used and is not counted.** The brief allowed one press of „Send self-bind
link" if no mail was waiting; a mail WAS waiting and it worked, so the bind cost zero interventions.

Secrets discipline for this step: the bind token, the passphrase and the box's root password each
live in a 0600 scratchpad file and were passed to `curl --data-urlencode name@file`, so no value
reached a command line. None of the three is written into this evidence, and the passphrase's only
description here is its shape (5 words, 35 characters).

### 0.2d — the box rotates its own root password at day-0, and my access died with it
At 20:18:58Z my SSH to the box still worked; at 20:19:04Z it answered „Permission denied", and the
background watcher lost access in the same window (its 20:19:27Z line came back empty). The bind at
20:18:15Z had started the day-0 install.

**This is the design, not a defect, and it was confirmed in source rather than guessed:**
`scripts/felhom-host-install.sh` (≈2079-2104) generates a strong `root@pam` password with `openssl
rand`, sets it through `chpasswd` on **stdin** (no argv, no log), and vaults it to the hub with
`PUT /api/v1/hosts/<id>/recovery-credential` over the enroll-authenticated channel. The password is
never logged, printed, or written to a file anywhere on the box. So the installer-time password I
typed into the Proxmox installer is dead by intent the moment day-0 runs.

**What this changes for tonight:** every accident that acts INSIDE the box (docker restart, tunnel
kill, filling the system disk) needs the hub-vaulted break-glass credential
(`/hosts/<host-id>/reveal-recovery-credential`), not the install password. Discovered at 22:19 CEST,
before the rounds began, rather than at 01:10 in the middle of round 5 — which is the only reason it
is a method note here and not an intervention later.

**How it was diagnosed honestly:** my first instinct was that I had broken my own environment. That
was ruled out first — the password file was still 21 bytes, the variable still 20 characters, and the
same credential had worked six seconds earlier. Only then was the box's own behaviour blamed, and
only after the installer source confirmed the mechanism.

### 0.2e — THE F-14 PATH, MEASURED LIVE FOR THE FIRST TIME
The brief named this as a claim that had never been measured: „the WG hook provisions by itself after
an acknowledged delete". Tonight it ran, unprompted, and the hub recorded it:

  Sep 16 20:18  appliance_bound (customer_selfbind)
      „Az ügyfél saját maga kötötte össze az új eszközt (bare-metal telepítés); a hozzáférést a
       doboz a következő lekérdezéskor megkapja."
  Sep 16 20:18  appliance_credential_delivered
      „Új eszköz (bare-metal telepítés) megkapta a hozzáférést és megkezdi a beállítást."
  Sep 16 20:18  claim_reissued_reenroll
      „Új beállító kódot küldtünk a szerver újratelepítése után (6. generáció) az ügyfél címére."
  Sep 16 20:19  offsite_reissued
      „Az offsite (házon kívüli) mentési hozzáférést újra kiadtuk — az új egyszeri jelszót a vezérlő
       a következő frissítéskor átveszi."
  Sep 16 20:19  **pbsdr_auto_reissue**
      „Previous key destroyed (acknowledged deletion) — credentials re-issued automatically."

That last line is the one that matters. Yesterday's box was removed through the **acknowledged**
delete flow, and tonight's box therefore got its off-site and PBS-DR credentials **with no operator
press at all** — exactly what the F-14 ruling of 2026-07-13 says should happen on that path, and the
half of that ruling nobody had yet watched happen.

**Consequence for this night's intervention count:** BOTH pre-declared presses are unnecessary.
O1 („Send self-bind link") was not needed because the automatic mail was waiting; O2 („Re-issue PBS
credentials") was not needed because the acknowledged-delete path re-issued by itself.
**Interventions so far: 0.**

Host enrolled as **`tester-1-022354`**, agent **0.131.0**, ONLINE, guests 0/0 at 20:20Z — the
customer guest is still being created from the golden.

### 0.2f — day-0 delivered tonight's golden, with no hand upgrade
  guest: 9201 „tester-1", running, created by the bootstrap from the golden
  controller image: gitea.dooplex.hu/admin/felhom-controller:**0.245.0** — „Up … (healthy)"
  agent: felhom-agent **0.131.0**
  host: `tester-1-022354`, ONLINE in the hub

The golden baked at 19:58Z tonight (sha 7a08aa1a…) is what this box installed. Nothing was upgraded
by hand, and the controller the customer will use is the release this drill is validating. That is
the delivery half of the chain: bake -> vouch -> a fresh box lands on it.

  both disks present to the box: `sda` 32 G (system, PVE + LVM) and **`sdb` 100 G** (the data disk,
  still unformatted — the household's drive, initialised through the storage page in the next step).

MY OWN MEASUREMENT ERROR, recorded: the first `lsblk` was piped through `head -12` and stopped one
line short of `sdb`. For a minute the box looked like it had NO data disk — a wrong answer produced
entirely by my own truncation, not by the box. Re-read without the pipe, `sdb 100G` is plainly there
and `qm config 336` still shows `scsi1` attached. An instrument that can cut off its own answer is
not a measurement.

### the dashboard setup code
The hub mailed a fresh „Új beállító kód — újratelepült a szervered" at **20:18:56Z** (the 6th
generation for this customer), 72 hours valid, delivered to `tester1@felhom.eu`. That is the code the
volunteer types on „A szerver beállítása" to set their own dashboard password — and it arrived by
itself, as part of the same automatic re-enrolment that needed no operator press.

### 0.2g — the dashboard claimed, by the volunteer, with the mailed code
The code from the 20:18:56Z mail („ősrégen-újraért-címbetű", 3 words, 72 h) was typed into
„A szerver beállítása" together with a 20-character password the household chooses.

  POST /claim            -> 302
  POST /login (new pw)   -> 302, and a `felhom_session` cookie was issued

**The second line is the proof; the first is only an attempt.** This repo has a standing trap that
an HTTP 200 (or a redirect) can be a refusal — the claim page re-rendering itself looks exactly like
success from the status code alone. The claim is called successful here because the password it set
then opened a session, which is the consequence a customer actually cares about.

Both values went in as FILES (`--data-urlencode name@file`, 0600, pushed with `pct push`), so
neither the setup code nor the new password ever reached a command line on the host or in the guest.

### 0.2h — tonight's release, seen working on a box that installed itself
The storage page of this fresh box carries `<form method="POST" action="/backup/escrow/banner/dismiss">`
— the **R-543 escrow reminder bar shipped in controller v0.245.0 a few hours ago**, rendering on a
box nobody had touched. It is there because the off-site tier was re-issued automatically at 20:19Z
and its escrow is not complete yet, which is precisely the state the bar exists for.

This is the first time that fix has been seen on a box that was not set up for the purpose of
testing it: the box installed itself from the published ISO, landed on tonight's golden, got its
off-site credentials with no press, and is now telling the household — on every page — that the
remote backup is paused until they create their recovery code. Creating it is the next step of the
guide, and of this drill.

### the data drive, as the box offers it
`/api/disks/candidates` reports exactly one initialisable device:
  /dev/sdb — 107 374 182 400 B (100 GB), QEMU HARDDISK, data_bearing=false, mountable=false
and separately the guest's own system volume as already mounted at /mnt/sys_drive. The empty
100 GB disk is the household's drive and the only thing offered for initialisation — the
data_bearing=false flag is the guard that keeps a drive with someone's files on it out of this list.

### Schedule: Phase 0 ran long, and the rounds shift with it — recorded, not quietly re-timed
The drawn schedule puts round 1 at 23:30 CEST. Phase 0 will not be finished by then: the box was
installed, bound, claimed and landed on tonight's golden without trouble, but working out how the
storage wizard actually initialises a disk took several rounds of discovery, because the wizard
submits through JavaScript (`POST /api/storage/init`) rather than a form, and I refused to guess the
endpoint after a guessed path cost a 403 and a wrong diagnosis in an earlier drill.

**What shifts and what does not.** The SCHEDULE — which action, on which app, under which accident,
in which order — is unchanged; it was drawn from the seed before anything ran and is fixed. Only the
wall-clock start moves, and the ~25-minute spacing is kept from the new start. The brief allows a
round to wait provided the wait is recorded; this is that record. The 05:00 stop rule is unchanged,
so a late start means the night may reach fewer than twelve rounds, and the morning verdict will say
how many actually ran rather than implying all twelve did.

### 0.2i — the household's drive, initialised through the wizard's own endpoint
The wizard submits by JavaScript, not by a form: `POST /api/storage/init` with
`{device, fstype, mount_name, label, set_default, confirmed, durable_id}` and the CSRF meta token,
then polls `GET /api/storage/init/status`. Both were read off the live page and confirmed in
`internal/web/storage_handlers.go:362` before anything was sent.

  POST /api/storage/init            -> 200  {"phase":"formatting","started":true}
  GET  /api/storage/init/status     -> phase **done**, started 20:27:21.79Z, updated 20:27:23.95Z
                                       where=/mnt/felhom-drives/hdd_1, error="" reason=""
  read back: /dev/sdb is ext4, durable_id `uuid:8f59ed90-c0e9-4e20-8584-d18af823c605`
             `df`: /dev/sdb 98G, 2.1M used, 93G free, mounted on /mnt/felhom-drives/hdd_1
             storage page: „hdd_1" labelled „Adatlemez", set as default

**The POST returning 200 is not the result** — it only says the job started. The format runs as a
background job precisely so a closed tab cannot abort it, so the phase poll is what says it worked,
and the `df` line is what says the household can use it.

**A known row met tonight, recorded rather than re-filed: R-542.** After the drive is formatted,
registered, mounted and made default, `/api/disks/candidates` STILL lists `/dev/sdb` under
`initialize` — now with `data_bearing:true, mountable:true` and its durable id. The same endpoint
also lists it under `attach`. That is exactly the behaviour R-542 describes (a registered, in-use
drive still offered under „initialize"), seen again on a fresh box.

### DECISION — the catalog bump was NOT pushed, and round 7 changes because of it
The bump was prepared (`templates/nextcloud/docker-compose.yml`: `redis:7-alpine` -> `redis:7.4-alpine`,
a tag checked to exist; `catalog_since` -> 2026-09-17) and then **reverted unpushed**, because the
catalog repo's own gate runner returned:

    image-resolvable     INCONCLUSIVE  (exit 2)
    volume-persistence   INCONCLUSIVE  (exit 2)
    UNDETERMINED (never a pass): image-resolvable, volume-persistence

with `volume-persistence` failing its OWN canary self-test —
    „ERROR: the prober failed its own canary — canary-clean: expected CLEAN, got UNDETERMINED …
     refusing to report a verdict: a broken detector reporting CLEAN is worse than no detector at all"

Neither inconclusive gate is caused by the one-line pin change: `catalog-since` and `engine-major`
both reported „0 compose file(s) changed" for the same range, i.e. they saw nothing of it. This is an
environment/prober problem, and it is **recorded, not worked around** — this repo's rule is that
UNDETERMINED is never a pass, and pushing past it would be exactly the habit the rule exists to stop.
Nothing was bypassed and `--no-verify` was not used; the tree was returned to clean.

**Consequence for the schedule, decided and recorded rather than improvised at 02:00:** round 7 drew
`update nextcloud`, and with no bump delivered there is no update for the guarded-update path to
apply. By the schedule's own constraint 3 — an `update` that cannot run becomes `use` — **round 7
runs as `use nextcloud` under its drawn accident (`internet gone 10min`)**, and the findings document
says so in that round's row. The guarded update is therefore NOT exercised tonight, and the morning
verdict must not claim it was.

### Reachability, settled before the rounds — and where the household loop actually runs
The customer guest is at **192.168.0.116**, on the LAN, and IS reachable from DooPlex:
    ping answers · ports 80 and 443 open · `Host: wiki.enkicsifelhom.hu` -> 301 (redirect to https)
So the `use` rounds can drive the apps' own front doors directly from the drill host, which is what
the brief asked for. Note the contrast with the box itself: the nested PVE (192.168.0.115) answers,
but the CONTROLLER only listens on its container address 172.17.0.2:8080 inside the guest — that is
why every dashboard action tonight is run as a script pushed into the guest.

**The background household loop runs ON THE VM (192.168.0.115), not on DooPlex** — a deviation from
the brief's „from the drill host", declared here. It was started before reachability from DooPlex had
been established, it is a `systemd-run` transient unit (`household.service`, active, logging to
/root/household.log), and moving a working measurement mid-night to gain nothing is a worse trade
than recording where it sits. It hits the guest's traefik from one hop closer than DooPlex would, so
it does NOT exercise the LAN path between DooPlex and the box; anything that breaks only on that hop
is invisible to it. The `use` rounds, which DO run from DooPlex, cover that hop.

### Fences and the pre-block baseline, checked mid-night (2026-09-16 ~21:38Z)
**Standing guests untouched.** `qm list` on demo-hp shows only VM **336** (this drill's own box);
containers 9201 and 9202 are running. 9201 still carries its full standing set —
adventurelog(+frontend,+postgres), **bentopdf**, bookstack(+db), calibre-web, cloudflared, docmost
(+postgres,+redis), felhom-controller, filebrowser, kimai(+db), opengist, paperless(+postgres,
+redis), privatebin, romm(+db,+redis), traefik — and 9202 its three (controller, filebrowser,
traefik). Nothing of theirs was stopped, removed or redeployed tonight. `drill-r50` does not exist on
this host and was not created.

**Firewall baseline on demo-hp, recorded BEFORE the internet-block rounds (7–9) need it:**
    iptables -S FORWARD   ->  `-P FORWARD ACCEPT`  (nothing else)
    physdev rules         ->  **0**
    net.bridge.bridge-nf-call-iptables = **0**

This is a control taken before the experiment, not after: rounds 7, 8 and 9 flip that sysctl to 1 and
insert two `--physdev-in` rules for this VM's tap. When they finish, the same three readings must
come back identical. Without the „before" reading, „it looks clean afterwards" would be an assertion
rather than a measurement — and this host also carries the two standing demo guests, so an abandoned
rule would be a fence breach rather than an untidy drill.
