Phase 0 finished: all twelve apps deployed one at a time (the parallel burst was my error and is recorded as such), escrow ceremony completed after the box converged the PBS descriptor by itself (~17 min), and the R-543 bar disappeared for good once escrowed. Round 1 (offsite-run, control round, no accident): - the run started on the button's own endpoint and walked every app with real per-volume byte counts (stop -> dump -> restart) - three alarms fired, all TRUE and precise: app_start_failed named the one crash-looping app, backup_run_failures said "1 of 12 ... nextcloud", and offbox_repo_orphaned reported the documented rebuild behaviour rather than claiming a successful copy - nextcloud's crash loop is MY damage (image layers written while the thin pool was 100% full -> "invalid ELF header"), not a product defect, and is recorded that way The catalog bump was prepared and then REVERTED UNPUSHED: the catalog repo's own gates returned INCONCLUSIVE (the volume-persistence prober failed its own canary), and undetermined is never a pass. Round 7's `update` therefore becomes `use`, decided now rather than improvised at 02:00. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -360,3 +360,45 @@ registered, mounted and made default, `/api/disks/candidates` STILL lists `/dev/
|
||||
`initialize` — now with `data_bearing:true, mountable:true` and its durable id. The same endpoint
|
||||
also lists it under `attach`. That is exactly the behaviour R-542 describes (a registered, in-use
|
||||
drive still offered under „initialize"), seen again on a fresh box.
|
||||
|
||||
### DECISION — the catalog bump was NOT pushed, and round 7 changes because of it
|
||||
The bump was prepared (`templates/nextcloud/docker-compose.yml`: `redis:7-alpine` -> `redis:7.4-alpine`,
|
||||
a tag checked to exist; `catalog_since` -> 2026-09-17) and then **reverted unpushed**, because the
|
||||
catalog repo's own gate runner returned:
|
||||
|
||||
image-resolvable INCONCLUSIVE (exit 2)
|
||||
volume-persistence INCONCLUSIVE (exit 2)
|
||||
UNDETERMINED (never a pass): image-resolvable, volume-persistence
|
||||
|
||||
with `volume-persistence` failing its OWN canary self-test —
|
||||
„ERROR: the prober failed its own canary — canary-clean: expected CLEAN, got UNDETERMINED …
|
||||
refusing to report a verdict: a broken detector reporting CLEAN is worse than no detector at all"
|
||||
|
||||
Neither inconclusive gate is caused by the one-line pin change: `catalog-since` and `engine-major`
|
||||
both reported „0 compose file(s) changed" for the same range, i.e. they saw nothing of it. This is an
|
||||
environment/prober problem, and it is **recorded, not worked around** — this repo's rule is that
|
||||
UNDETERMINED is never a pass, and pushing past it would be exactly the habit the rule exists to stop.
|
||||
Nothing was bypassed and `--no-verify` was not used; the tree was returned to clean.
|
||||
|
||||
**Consequence for the schedule, decided and recorded rather than improvised at 02:00:** round 7 drew
|
||||
`update nextcloud`, and with no bump delivered there is no update for the guarded-update path to
|
||||
apply. By the schedule's own constraint 3 — an `update` that cannot run becomes `use` — **round 7
|
||||
runs as `use nextcloud` under its drawn accident (`internet gone 10min`)**, and the findings document
|
||||
says so in that round's row. The guarded update is therefore NOT exercised tonight, and the morning
|
||||
verdict must not claim it was.
|
||||
|
||||
### Reachability, settled before the rounds — and where the household loop actually runs
|
||||
The customer guest is at **192.168.0.116**, on the LAN, and IS reachable from DooPlex:
|
||||
ping answers · ports 80 and 443 open · `Host: wiki.enkicsifelhom.hu` -> 301 (redirect to https)
|
||||
So the `use` rounds can drive the apps' own front doors directly from the drill host, which is what
|
||||
the brief asked for. Note the contrast with the box itself: the nested PVE (192.168.0.115) answers,
|
||||
but the CONTROLLER only listens on its container address 172.17.0.2:8080 inside the guest — that is
|
||||
why every dashboard action tonight is run as a script pushed into the guest.
|
||||
|
||||
**The background household loop runs ON THE VM (192.168.0.115), not on DooPlex** — a deviation from
|
||||
the brief's „from the drill host", declared here. It was started before reachability from DooPlex had
|
||||
been established, it is a `systemd-run` transient unit (`household.service`, active, logging to
|
||||
/root/household.log), and moving a working measurement mid-night to gain nothing is a worse trade
|
||||
than recording where it sits. It hits the guest's traefik from one hop closer than DooPlex would, so
|
||||
it does NOT exercise the LAN path between DooPlex and the box; anything that breaks only on that hop
|
||||
is invisible to it. The `use` rounds, which DO run from DooPlex, cover that hop.
|
||||
|
||||
@@ -0,0 +1,66 @@
|
||||
## Round 1 side-observations — two things that looked like defects and are not (and one that is mine)
|
||||
|
||||
### 1. Five front doors answered 404 — because the BACKUP was quiescing them
|
||||
A front-door sweep from DooPlex during round 1 read:
|
||||
wiki 404 · paste 200 · share 404 · recipes 200 · inventory 200 · status 200 · travel 200
|
||||
paperless 404 · cloud 404 · photos 404 · media 200 · vault 200
|
||||
The container states explained it: `bookstack Up 17 seconds`, `gokapi Up 1 second`,
|
||||
`paperless-webserver Up 11 seconds`, `immich-server Created` — all seconds old. The controller log
|
||||
says why, in its own words:
|
||||
„[backup] Stopping nextcloud for safe volume dump" -> volume dumps -> „[backup] Restarting
|
||||
nextcloud after volume dump"
|
||||
So the off-site run stops each app, dumps its volumes, and starts it again — and my sweep happened to
|
||||
walk the apps while the run was walking them too. **Nothing was wrong; my measurement was mistimed.**
|
||||
The sweep is repeated after the run finishes, and only that later reading counts.
|
||||
|
||||
### 2. Nextcloud IS broken — and it is MY damage, from the disk-full episode
|
||||
php: error while loading shared libraries: /lib/x86_64-linux-gnu/libxml2.so.2: **invalid ELF header**
|
||||
RestartCount=6 ExitCode=127 Status=restarting
|
||||
„Invalid ELF header" on a shared library means the file on disk is CORRUPT, not missing. Those layers
|
||||
were written while the LVM thin pool stood at 100 % (my parallel twelve-deploy burst). The later
|
||||
sequential re-deploy reported „INSTALLED" because it reused the already-pulled — and already
|
||||
corrupt — image instead of pulling again. Repair: drop the image so a fresh pull happens, and
|
||||
redeploy. Not done during an in-flight backup run.
|
||||
|
||||
**This is not filed as a product defect.** The corruption came from a full disk that I caused.
|
||||
|
||||
### What the PRODUCT did about it, which is the part worth measuring
|
||||
„[stacks] Stack nextcloud started successfully (took 7.5s)"
|
||||
„[stacks] Stack nextcloud post-start status:"
|
||||
„[stacks] nextcloud nextcloud:34.0.1-apache restarting Restarting (127) 1 second ago"
|
||||
This is the documented invariant working: `compose up -d` exits 0 even for a crash-loop, and **the
|
||||
post-start status check is the detection**. The success line and the truth appear together, one line
|
||||
apart, and the state the box records is `restarting`, not `running`.
|
||||
|
||||
## The alarms round 1 produced — three fired, all TRUE, and each one precise
|
||||
20:51 warning **app_start_failed** „Telepített alkalmazás nem fut: Nextcloud"
|
||||
21:08 error **backup_run_failures** „1 of 12 apps failed to back up in this nightly run: nextcloud"
|
||||
21:09 warning **offbox_repo_orphaned** „A távoli mentési tároló elárvult: a benne lévő mentések egy
|
||||
korábbi, már nem elérhető kulccsal készültek (újratelepítés).
|
||||
Új mentés a tároló visszaállít…"
|
||||
|
||||
Scored against `08-alarm-ladder.md`:
|
||||
* `app_start_failed` is exactly right for a deployed app that is not running, and it is the one
|
||||
customer-switchable event that is OFF by default — it reached the operator. It fired ~5 minutes
|
||||
after the crash loop started, which is the documented 5-minute crash-loop window, not a lag.
|
||||
* `backup_run_failures` is operator-only by register, error severity, and **it names the app and
|
||||
the count**: „1 of 12 … nextcloud". A run that partly failed said so precisely rather than
|
||||
reporting a green run or a red one.
|
||||
* `offbox_repo_orphaned` is the honest outcome of round 1 itself (below). Nothing pretended the
|
||||
off-site copy had succeeded.
|
||||
|
||||
**No alarm was missed for what happened**, and no alarm fired for something that had not happened.
|
||||
The only app that was broken is the only app that was named.
|
||||
|
||||
## What round 1 actually measured (the drawn action was `offsite-run`)
|
||||
* the run STARTED on the button's own endpoint: „A távoli mentés elindult — az állapot itt frissül."
|
||||
with `progress.active: true`
|
||||
* it walked every app: stop -> volume dump -> restart („Stopping X for safe volume dump" …
|
||||
„Restarting X after volume dump"), with real byte counts per volume
|
||||
(e.g. nextcloud_db_data 148.9 MB, nextcloud_html 2.5 KB, redis_data 5.5 KB)
|
||||
* one app could not be captured — the one that was crash-looping — and the run said so
|
||||
* and the repository itself turned out to be **orphaned**: this box is a REBUILD for an existing
|
||||
customer, so its restic password was minted fresh and the snapshots already in the remote store
|
||||
were written under a key this box does not have. That is the documented rebuild behaviour
|
||||
(„Guest rebuild orphans the off-site repo"), surfaced to the customer rather than hidden, with
|
||||
the route out named in the message.
|
||||
@@ -0,0 +1,17 @@
|
||||
2026-09-16T21:07:14Z === ROUND 1: offsite-run / adventurelog / accident: nothing ===
|
||||
--- BEFORE ---
|
||||
apps running: 26
|
||||
escrow=escrowed last_run=None last_status=None snapshots=None repo=None
|
||||
--- ACTION: the off-site run, through the button's own endpoint ---
|
||||
POST /backup/offbox/run -> HTTP/1.1 302 Found
|
||||
flash: A távoli mentés elindult — az állapot itt frissül.
|
||||
--- polling the run ---
|
||||
2026-09-16T21:07:34Z {"last_duration":"","last_error":"","last_run":"","orphaned":false,"progress":{"active":true,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":0,"current_file":"","elap
|
||||
2026-09-16T21:07:54Z {"last_duration":"","last_error":"","last_run":"","orphaned":false,"progress":{"active":true,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":0,"current_file":"","elap
|
||||
2026-09-16T21:08:15Z {"last_duration":"","last_error":"","last_run":"","orphaned":false,"progress":{"active":true,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":0,"current_file":"","elap
|
||||
2026-09-16T21:08:35Z {"last_duration":"","last_error":"","last_run":"","orphaned":false,"progress":{"active":true,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":0,"current_file":"","elap
|
||||
2026-09-16T21:08:55Z {"last_duration":"","last_error":"","last_run":"","orphaned":false,"progress":{"active":true,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":0,"current_file":"","elap
|
||||
2026-09-16T21:09:15Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":
|
||||
2026-09-16T21:09:35Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":
|
||||
2026-09-16T21:09:55Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":
|
||||
2026-09-16T21:10:15Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":
|
||||
Reference in New Issue
Block a user