## Round 1 side-observations — two things that looked like defects and are not (and one that is mine)

### 1. Five front doors answered 404 — because the BACKUP was quiescing them
A front-door sweep from DooPlex during round 1 read:
    wiki 404 · paste 200 · share 404 · recipes 200 · inventory 200 · status 200 · travel 200
    paperless 404 · cloud 404 · photos 404 · media 200 · vault 200
The container states explained it: `bookstack Up 17 seconds`, `gokapi Up 1 second`,
`paperless-webserver Up 11 seconds`, `immich-server Created` — all seconds old. The controller log
says why, in its own words:
    „[backup] Stopping nextcloud for safe volume dump" -> volume dumps -> „[backup] Restarting
     nextcloud after volume dump"
So the off-site run stops each app, dumps its volumes, and starts it again — and my sweep happened to
walk the apps while the run was walking them too. **Nothing was wrong; my measurement was mistimed.**
The sweep is repeated after the run finishes, and only that later reading counts.

### 2. Nextcloud IS broken — and it is MY damage, from the disk-full episode
    php: error while loading shared libraries: /lib/x86_64-linux-gnu/libxml2.so.2: **invalid ELF header**
    RestartCount=6  ExitCode=127  Status=restarting
„Invalid ELF header" on a shared library means the file on disk is CORRUPT, not missing. Those layers
were written while the LVM thin pool stood at 100 % (my parallel twelve-deploy burst). The later
sequential re-deploy reported „INSTALLED" because it reused the already-pulled — and already
corrupt — image instead of pulling again. Repair: drop the image so a fresh pull happens, and
redeploy. Not done during an in-flight backup run.

**This is not filed as a product defect.** The corruption came from a full disk that I caused.

### What the PRODUCT did about it, which is the part worth measuring
    „[stacks] Stack nextcloud started successfully (took 7.5s)"
    „[stacks] Stack nextcloud post-start status:"
    „[stacks]   nextcloud   nextcloud:34.0.1-apache   restarting   Restarting (127) 1 second ago"
This is the documented invariant working: `compose up -d` exits 0 even for a crash-loop, and **the
post-start status check is the detection**. The success line and the truth appear together, one line
apart, and the state the box records is `restarting`, not `running`.

## The alarms round 1 produced — three fired, all TRUE, and each one precise
    20:51  warning  **app_start_failed**      „Telepített alkalmazás nem fut: Nextcloud"
    21:08  error    **backup_run_failures**   „1 of 12 apps failed to back up in this nightly run: nextcloud"
    21:09  warning  **offbox_repo_orphaned**  „A távoli mentési tároló elárvult: a benne lévő mentések egy
                                               korábbi, már nem elérhető kulccsal készültek (újratelepítés).
                                               Új mentés a tároló visszaállít…"

Scored against `08-alarm-ladder.md`:
  * `app_start_failed` is exactly right for a deployed app that is not running, and it is the one
    customer-switchable event that is OFF by default — it reached the operator. It fired ~5 minutes
    after the crash loop started, which is the documented 5-minute crash-loop window, not a lag.
  * `backup_run_failures` is operator-only by register, error severity, and **it names the app and
    the count**: „1 of 12 … nextcloud". A run that partly failed said so precisely rather than
    reporting a green run or a red one.
  * `offbox_repo_orphaned` is the honest outcome of round 1 itself (below). Nothing pretended the
    off-site copy had succeeded.

**No alarm was missed for what happened**, and no alarm fired for something that had not happened.
The only app that was broken is the only app that was named.

## What round 1 actually measured (the drawn action was `offsite-run`)
  * the run STARTED on the button's own endpoint: „A távoli mentés elindult — az állapot itt frissül."
    with `progress.active: true`
  * it walked every app: stop -> volume dump -> restart („Stopping X for safe volume dump" …
    „Restarting X after volume dump"), with real byte counts per volume
    (e.g. nextcloud_db_data 148.9 MB, nextcloud_html 2.5 KB, redis_data 5.5 KB)
  * one app could not be captured — the one that was crash-looping — and the run said so
  * and the repository itself turned out to be **orphaned**: this box is a REBUILD for an existing
    customer, so its restic password was minted fresh and the snapshots already in the remote store
    were written under a key this box does not have. That is the documented rebuild behaviour
    („Guest rebuild orphans the off-site repo"), surfaced to the customer rather than hidden, with
    the route out named in the message.

## The 404s, settled with controls — and my first explanation was WRONG
I first explained the five 404s as „my sweep ran while the backup was quiescing apps". After the
backup finished, four of the five were **still** 404, so that explanation was wrong for them. What it
actually is, established with both controls rather than by inference:

    negative control  nosuchapp.enkicsifelhom.hu -> 404, **19 bytes**, „404 page not found"
    wiki / share / paperless / photos            -> 404, **19 bytes**, „404 page not found"  (identical)
    positive control  status.enkicsifelhom.hu    -> **200, 1200 bytes**

Identical to the no-such-host control, byte for byte: **traefik has no route for them.** The apps did
not answer 404 — nothing answered. And the reason there is no route is the container state, because
traefik only routes to a container that is up and healthy:

    bookstack             Up 5 minutes (**unhealthy**)
    gokapi                **Restarting (0)**
    immich-server         **Restarting (1)**
    paperless-webserver   Up 9 seconds (health: starting)   <- this one was simply still starting
    nextcloud             Up 20 seconds (health: starting)  <- the image re-pull REPAIRED it

So the sequence is: my disk-full episode corrupted image layers -> containers crash-loop or stay
unhealthy -> traefik has no route -> the front door 404s. Nextcloud proves the chain from the other
end: drop the corrupt image, let compose pull it again, and the app comes up. The same repair is
applied to bookstack, gokapi and immich, with each one's failing log lines captured BEFORE the fix so
the cause is recorded and not just the cure.

**None of this is filed as a product defect** — the corruption is mine. What the product did with it
was correct throughout: it refused to route to unhealthy containers, it reported `app_start_failed`
for the app that was down, and it named that same app as the one that failed to back up.

## The nextcloud repair, with its own record (not inferred from a health badge)
    stop  through the product's API            -> 200
    docker image rm nextcloud:34.0.1-apache    -> Deleted: sha256:f2ec4de84f559f5c7be4233b589cdbdbb5507807e05621b77320edd55a1f2a0f
    start through the product's API            -> 200  (compose re-pulls the missing image)
    t+15s  nextcloud Up 15 seconds (health: starting)

That closes the chain from both ends: corrupt layer -> `invalid ELF header` -> exit 127 crash loop
-> no traefik route -> 404 at the front door; and then drop the image -> fresh pull -> the app comes
up. The repair was done through the product's own stop/start endpoints, not by hand-running compose,
so the controller's record of the app stayed consistent throughout.

## The cause evidence contradicts my own hypothesis — recorded before the repair overwrote it
I had assumed all the broken apps were corrupt-layer damage like nextcloud. The logs captured BEFORE
touching them say otherwise:

  * **bookstack** — a perfectly normal startup: „Waiting for DB to be available" … „INFO Nothing to
    migrate." … „[custom-init] No custom files found, skipping…" … „[ls.io-init] done." No error at
    all. It was merely marked `unhealthy`, i.e. its healthcheck was not passing — which is enough for
    traefik to refuse it a route, but is NOT corruption. Dropping its image was probably unnecessary.
  * **gokapi** — **no log output whatsoever**, and `Restarting (0)`: it exits cleanly and instantly.
    That is a configuration/first-run shape, not a corrupt binary.
  * **immich-server** — `file: 'auth.c', line: '331', routine: 'auth_failed'` under Node.js. That is
    **PostgreSQL refusing the password**, not a broken library.

### The likely real cause, and it is mine: regenerated secrets over existing databases
My sequential re-seed called `gen` again for every app, so each one was deployed with **freshly
generated DB passwords** — while the database volumes created during the first (failed, parallel)
burst still hold databases initialised with the FIRST passwords. Postgres/MariaDB keep the password
from `POSTGRES_PASSWORD`/`MYSQL_PASSWORD` at volume-init time and ignore it afterwards, so the app
then authenticates with a password the database has never heard of. `auth_failed` is exactly that.

**Consequence:** re-pulling an image cannot fix this, and neither can restarting. The apps whose
volumes predate the re-seed need a clean removal (with their data) and a fresh deploy — which costs
nothing here because none of them holds real household data yet.

**Not a product defect.** The product deployed exactly what I asked it to deploy, twice, with the
values I supplied each time.

## CORRECTION — „nextcloud is repaired" was WRONG, and the container badge is why I believed it
I wrote that the image re-pull repaired nextcloud, on the strength of `Up (healthy)`. Checking the
log with timestamps against the guest's own clock:

    guest time now                     2026-09-16T21:16:02Z
    2026-09-16T21:14:09Z  „Access denied for user 'nextcloud'@'172.21.0.4' (using password: YES)"
    2026-09-16T21:14:19Z  „Error while trying to create admin account: … Access denied for user
                            'nextcloud'@'172.21.0.4' (using password: YES)"

Two minutes old — **current, not leftovers from the crash-looping period.** The re-pull did fix the
corrupt library (the app now starts instead of exiting 127), but the SECOND fault is still there: the
database was initialised with the first deploy's password and the app now presents the re-seed's new
one. It could not even create its admin account.

**And the container says `healthy` the whole time.** The healthcheck is the app image's own — Apache
answers, so the probe passes — while the application cannot reach its database at all. This is the
project's standing lesson in another costume: **a health signal is not a data signal**, and „healthy"
answered a different question from the one I was asking.

So nextcloud joins gokapi and paperless in the clean remove-and-redeploy. Retracted here rather than
left standing, because the wrong version of this paragraph was already written two entries above.

## An HTTP 000 is „no answer", not „it failed" — and the action still happened
The immich repair's start call returned:
    start -> 000
`000` is curl's code for „no response received", not a refusal by the server. Twenty seconds later:
    immich-server  Up 4 seconds (health: starting)
So the start DID take effect; only the answer was lost (the controller was busy stopping/starting
several stacks at once). Recorded because reading `000` as „the start failed" would have led me to
issue a second start on an app that was already coming up.

## Household state at 21:17Z, and the decision to move round 2
    immich-server        Up 4 seconds (health: starting)   — DB password mismatch still expected
    gokapi               Restarting (1)                    — clean remove+redeploy in progress
    bookstack            Up about a minute (**unhealthy**) — healthcheck not passing
    nextcloud            Up about a minute (healthy)       — but DB-broken, see the correction above
    paperless-webserver  Up 8 seconds                      — clean remove+redeploy in progress
    (the other seven apps: up and serving)

**Decision: round 2 moves to ~23:45 CEST.** Five of twelve apps are mid-repair from damage I caused,
and a chaos round measures nothing if the household is already broken before the accident lands. The
SCHEDULE is unchanged — same actions, same apps, same accidents, same order, drawn from the seed
before anything ran; only the wall clock moves, and the 05:00 stop is unchanged. At ~25-minute
spacing a 23:45 start still reaches round 12 by about 04:20.

The queue behind the current job: **nextcloud, immich and bookstack** get the same clean
remove-and-redeploy. They are done one at a time, never two repair jobs at once on the same box —
concurrent stop/start on one controller is what produced half of tonight's confusion already.

## The diagnosis is confirmed by the cure
gokapi had already had a fresh image pulled and was STILL `Restarting (1)`. After a clean removal
(with data) and a fresh deploy with one consistent set of secrets:

    gokapi               **Up 47 seconds (healthy)**
    paperless-ngx        INSTALLED at 2026-09-16T21:17:15Z
    paperless-postgres   Up 19 seconds (healthy)
    paperless-redis      Up 19 seconds (healthy)

A new image did not fix it and a new database did — which is the difference between the corrupt-layer
fault (nextcloud's exit 127) and the regenerated-password fault (everything else). Two faults, both
mine, with different cures, and it took separating them to fix either.

Household loop through all of this: unit active, **4 sampled apps since the classifier fix, 0
failures** — the box kept serving its other apps while five were being torn down and rebuilt.

## A LATE GHOST FROM THIS ROUND, recorded 2026-09-17T01:26 CEST (23:26Z)
Round 1's off-site poll loop reported "completed, exit code 0" TWO HOURS after it stopped working.
  last real output line: 2026-09-16T21:24:56Z
  task actually ended:   2026-09-16T23:24:56Z, with
      "Read from remote host 192.168.0.115: Connection reset by peer"
      "client_loop: send disconnect: Broken pipe"
It went silent during round 2's power cut (the box it was polling was abruptly stopped), and the
SSH session then hung, producing nothing, until TCP reset it two hours later.

Three things worth keeping, all of them this project's recurring classes:
  1. A HUNG connection and a FINISHED one look identical from outside: both produce no new output.
     Silence is not completion. The loop's own timestamps are what distinguish them, which is why
     every poll line carries one.
  2. The task EXITED 0 while its connection had been reset. Another instance of "exit codes that
     lie" - the exit status described the local shell, not the remote work.
  3. A completion notice arriving two hours late could easily be read as a FRESH result for
     whatever round is running now. It was checked against its own content before being believed.
It did nothing to the box after 21:24:56Z, so no round is contaminated. Recorded, not hidden.
