Files
felhom.eu/documentation/audits/update-night-2026-09-21/bad-days/B5-safety-dump/result.json
T
admin 9c69b3ff07
gates / gates (push) Successful in 27s
Update night: Phases 2-4 evidence — both engines, the unattended HOLD, and five new findings
Evidence off the machine at the end of the phases that produced it (R-320). Teardown follows.

PHASE 2 — the two database engines, through the REAL Update button:
- MariaDB 11.6 -> 12.3 on nextcloud: PROVEN, and pressed through the button for the first time.
  All four SPIKE-r459 observables: the datadir's own record moved 11.6.2 -> 12.3.3; the engine
  itself says "already upgraded ... no need to run mariadb-upgrade again"; the entrypoint says
  "Major version upgrade detected ... Check required!" and then STARTED and FINISHED it (not the
  `skipped due to $MARIADB_AUTO_UPGRADE` line R-459 feared); and the engine took its own
  pre-upgrade backup, 631 905 B. The seeded Nextcloud account read back.
- PostgreSQL 16 -> 17 on docmost: FAILED exactly as R-463 predicted and nobody had measured.
  5.1 s to held; the pin named 17 while nothing ran; the restore brought it back in 29.1 s.
  The engine's REFUSAL LINE was destroyed by failAndHold before any probe could read it, so it
  was REPRODUCED INDEPENDENTLY with a control on every step (R-320).

PHASE 3 — the bad days. B1 produced THE UNATTENDED HOLD, which this project has never had: the
caller pressed once with nobody watching, the app held after 312.9 s, and passes 2 and 3 pressed
nothing. B2 put the pin back on a pull failure in 1.0 s. B3 refused `busy` six times. B4 showed
there is NO single-flight — 5 of 5 updates ran at once and all ended honest. B5 cut the power in
`backing-up` and the box recovered itself and said so. B7 refused under the 2 GB floor. B9 found
R-458's risk narrower than the row states.

PHASE 4 — every badge on the box is TRUE, and the held app answers all four of Q4's questions.

FINDINGS, five new and three corrections to existing rows. The one that matters: R-618 is P1 —
two templates name a health probe the app does not answer, and because the guarded update waits
on that same probe, a SUCCESSFUL update ends by STOPPING a working app. Measured: tandoor served
HTTP 200 on the new version at four samples across five minutes and was then stopped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 22:13:57 +02:00

43 lines
2.4 KiB
JSON

{
"leg": "B5-safety-dump",
"app": "bentopdf",
"phase_targeted": "safety-dump",
"observables_before": {
"pinned_images": {
"bentopdf": "localhost:5000/drill/pdf:2.0.1"
},
"installed_images": {
"bentopdf": "localhost:5000/drill/pdf:2.0.1"
},
"catalog_images": {
"bentopdf": "localhost:5000/drill/pdf:2.0.1"
},
"live_compose_image_lines": [
"image: localhost:5000/drill/pdf:2.0.1"
],
"docker_inspect": [
"bentopdf localhost:5000/drill/pdf:2.0.1 running=true restarts=0"
]
},
"backup_artefacts_before": "=== the app's own recovery unit + its db dumps, with sizes and times\n2026-09-21 20:04 286 /mnt/sys_drive/felhom-data/backups/primary/bentopdf/compose/app.yaml\n2026-09-21 20:04 1006 /mnt/sys_drive/felhom-data/backups/primary/bentopdf/manifest.json\n2026-09-21 20:04 1106 /mnt/sys_drive/felhom-data/backups/primary/bentopdf/compose/docker-compose.yml\n2026-09-21 20:04 2961 /mnt/sys_drive/felhom-data/backups/primary/bentopdf/compose/.felhom.yml\n=== any temp/partial names left behind\n=== the unit manifest, if there is one\n--- /mnt/sys_drive/felhom-data/backups/primary/bentopdf/manifest.json\n{\n \"schema_version\": 2,\n \"app_name\": \"bentopdf\",\n \"display_name\": \"BentoPDF\",\n \"controller_version\": \"0.261.0\",\n \"created_at\": \"2026-09-21T20:04:43Z\",\n \"drive\": \"/mnt/sys_drive\",\n \"namespace_root\": \"/mnt/sys_drive/felhom-data\",\n \"image_pins\": [\n \"localhost:5000/drill/pdf:2.0.1\"\n ],\n \"secret_env_vars\": null,\n \"data_key_env_vars\": null,\n \"secret_source\": \"portable secrets (data keys, DB passwords, internal signing secrets) are IN this unit's compose/app.yaml (0600); internet-reachable admin logins are NOT, and come from the guest's app.yaml or are regenerated on restore\",\n \"config_files\": [\n \"docker-compose.yml\",\n \".felhom.yml\",\n \"app.yaml\"\n ],\n \"db_dumps\": [],\n \"vo\n=== ZERO-BYTE files under this app's backups (a half-write that still looks like a file)\n",
"cut": {
"pressed": true,
"phases_seen": [
{
"t": 0.022,
"phase": "backing-up"
},
{
"t": 0.248,
"phase": "pulling"
},
{
"t": 0.473,
"phase": "failed"
}
],
"cut_decided_at": null,
"missed": "safety-dump",
"ended_phase": "failed"
}
}