c85262111c
gates / gates (push) Successful in 23s
Part 0 — floor raised to 0.260.0, MinAgent 0.131.0 declared. 3 boxes below, all down or blocked; both demo boxes SERVED. Part 1 (R-610) — the DANGEROUS power cut, measured three times with three apps and two cut mechanisms. All ended honest: resumed, completed, and pinned/installed/live compose/docker inspect all agreed. vikunja's 2.6.0 migration had ALREADY run 0.64 s after the cut decision and the seeded data read back intact — so the branch that is one step from old-binary-on-migrated-database is now evidence, not argument. Instrument limit stated: `starting` lasts under a second; all three landed in `verifying`, which RecoverUpdates handles in the same branch. Part 3 (R-611) — the night the previous session skipped without saying so. An app updated with nobody pressing anything; a terminally-refused app was pressed exactly once and never again over three passes. The unattended HOLD was NOT produced: the within-a-major rule correctly refused the broken edge before it was attempted, so Q4 still rests on the attended hold from slice 4. Said plainly rather than implied. Rows: closed R-608/609/610/611; opened R-612 (P1 wishlist unusable on a fresh install, and its error is a lie), R-613 (uptime-kuma healthy on its setup wizard), R-614 (stale update phase survives a redeploy). R-520's pointer corrected. Catalog: two drill pairs, both reverted; every image line byte-identical to ff9717d3. The alpine:3.20 negative control a security review flagged is cleared. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
129 lines
8.3 KiB
Plaintext
129 lines
8.3 KiB
Plaintext
# 05 — SCENARIO B: uptime-kuma 2.4.0 -> 2.5.0, cut during `verifying` (the health wait)
|
|
# guest 9202 (demo-hp-scratch), controller 0.260.0, 2026-09-21
|
|
#
|
|
# VERDICT: the box ended HONEST. The update resumed after the reboot and finished `done`;
|
|
# all four version observables agree on 2.5.0; the seeded monitor read back through
|
|
# uptime-kuma's own socket.io front door; no journal survived; the household sentence is
|
|
# „Naprakesz" / "Up to date".
|
|
#
|
|
# INSTRUMENT FINDING (this one matters for reading BOTH A and B):
|
|
# `pct stop 9202` RETURNS after ~3.0-3.8 s, but the guest stops answering after ~1.1 s.
|
|
# Measured here by firing the cut in a THREAD and continuing to poll through it:
|
|
# decision 12:32:32.271, last successful API sample 12:32:33.345 (+1 075 ms),
|
|
# command return 12:32:35.303 (+3 032 ms).
|
|
# So the kill lands early and the rest of the 3 s is teardown. The practical consequence,
|
|
# proven in Scenario A: `starting` lasts about 0.6 s on this box, so NO cut driven this way
|
|
# can land inside `starting` — by the time the guest dies the update is in `verifying`.
|
|
# A and B therefore exercise the SAME RecoverUpdates branch
|
|
# (update.go:909 `case UpdatePhaseStarting, UpdatePhaseVerifying`). Stated, not hidden.
|
|
|
|
## (1) TIMESTAMP TABLE
|
|
poll target : GET /api/stacks/uptime-kuma, HTTPS keep-alive from DooPlex
|
|
measured poll cadence : ~21 ms (per-sample HTTP cost 1.1-1.3 ms + 20 ms sleep)
|
|
Update pressed : 12:32:26.374 UTC (POST /api/stacks/uptime-kuma/update -> accepted)
|
|
phase AT THE DECISION : "verifying" observed 12:32:32.271 UTC
|
|
cut command : ssh -S <prewarmed master> demo-hp 'pct stop 9202'
|
|
cut command LATENCY : 3 016 ms (rc=0, returned 12:32:35.303 UTC, no stdout/stderr)
|
|
box stopped ANSWERING : 12:32:33.345 UTC was the LAST successful sample (+1 075 ms)
|
|
phase the box DIED in : "verifying" (update-journal.json read off the stopped guest)
|
|
guest restarted : pct start 9202; controller up 12:33:21 UTC
|
|
update concluded : 12:33:26 UTC, "DONE in 1m0s"
|
|
|
|
## (1b) the phase trace, verbatim from the poller
|
|
12:32:25.147 updating=False phase='' state=running err='' hold=''
|
|
12:32:26.373 updating=True phase='checking' state=running err='' hold=''
|
|
12:32:26.394 updating=True phase='safety-dump' state=running err='' hold=''
|
|
12:32:26.416 updating=True phase='pulling' state=running err='' hold=''
|
|
12:32:27.627 updating=True phase='starting' state=running err='' hold=''
|
|
12:32:28.748 updating=True phase='starting' state=degraded err='' hold=''
|
|
12:32:32.271 updating=True phase='verifying' state=starting err='' hold=''
|
|
--- DECISION at 12:32:32.271: phase='verifying' -> firing (in a thread): ssh -S /tmp/cm-demohp demo-hp 'pct stop 9202'
|
|
12:32:33.367 ERR ConnectionResetError: [Errno 104] Connection reset by peer
|
|
12:32:33.393 ERR TimeoutError: timed out
|
|
|
|
=== TIMESTAMP TABLE (uptime-kuma) ===
|
|
phase at decision : 'verifying' at 12:32:32.271 UTC
|
|
cut command : ssh -S /tmp/cm-demohp demo-hp 'pct stop 9202'
|
|
cut command latency : 3016 ms (rc=0, returned 12:32:35.303 UTC)
|
|
cut stdout/stderr : '' / ''
|
|
LAST SUCCESSFUL SAMPLE : 12:32:33.345 UTC <- the box was still answering here
|
|
i.e. 1075 ms after the decision
|
|
|
|
## (1c) update-journal.json read from the STOPPED guest (pct mount 9202 first; pct unmount after)
|
|
POSITIVE CONTROL: catalog-cache present -> real dir
|
|
--- journal ---
|
|
{
|
|
"updates": {
|
|
"uptime-kuma": {
|
|
"phase": "verifying",
|
|
"started_at": "2026-09-21T12:32:26.364656602Z",
|
|
"prev_pin": {
|
|
"uptime-kuma": "louislam/uptime-kuma:2.4.0"
|
|
},
|
|
"prev_compose": "/opt/docker/stacks/uptime-kuma/pre-update-compose.yml",
|
|
"prev_applied": "/opt/docker/stacks/uptime-kuma/pre-update-applied.yml",
|
|
"proven_copy_at": "2026-09-21T12:16:12Z",
|
|
"proven_tier": 1
|
|
}
|
|
}
|
|
}
|
|
## (2) RecoverUpdates log lines after the restart, VERBATIM
|
|
2026/09/21 12:33:21 update.go:909: [WARN] [stacks] update recovery: uptime-kuma was interrupted in verifying (started 2026-09-21T12:32:26Z) — the new version may have run; marking it Updating and RESUMING the health wait
|
|
2026/09/21 12:33:21 update.go:951: [INFO] [stacks] update uptime-kuma: resuming after a controller restart — `up -d` then the health wait
|
|
2026/09/21 12:33:21 update.go:858: [INFO] [stacks] update uptime-kuma: phase verifying
|
|
2026/09/21 12:33:26 update.go:652: [INFO] [stacks] update uptime-kuma: healthy after 5s (the app's health check passed)
|
|
2026/09/21 12:33:26 update.go:658: [INFO] [stacks] update uptime-kuma: DONE in 1m0s
|
|
|
|
## (3) THE FOUR VERSION OBSERVABLES, SIDE BY SIDE (after recovery)
|
|
1_pinned_images : {"uptime-kuma": "louislam/uptime-kuma:2.5.0"}
|
|
2_installed_images : {"uptime-kuma": "louislam/uptime-kuma:2.5.0"} | digest: {"uptime-kuma": "sha256:a8610b3b4c38077922b"}
|
|
(updating=False phase=done label=Frissítve err=- hold=-)
|
|
3_live compose line: image: louislam/uptime-kuma:2.5.0
|
|
4_docker inspect : louislam/uptime-kuma:2.5.0 | RepoDigest: sha256:a8610b3b4c38077922b | started 2026-09-21T12:33:19.417787988Z
|
|
-> all four name louislam/uptime-kuma:2.5.0; installed digest and the running
|
|
container's RepoDigest are the same sha256:a8610b3b4c38077922b...
|
|
|
|
## (4) the seeded data, read back through uptime-kuma's OWN front door (socket.io, the same
|
|
## API the browser uses), run from inside the container with its own socket.io-client
|
|
MONITORS: ["SEED-KUMA-CANARY-4a91c7-20260921"]
|
|
POSITIVE CONTROL: seed present = true
|
|
NEGATIVE CONTROL: absent name present = false
|
|
read-back finished
|
|
NOTE: the FIRST read-back attempt, run 12 s after the app came up, TIMED OUT waiting for
|
|
the monitorList event — the app was listening but not yet serving. Re-run 20 s later it
|
|
answered. Recorded because a single timeout here reads exactly like data loss and is not.
|
|
|
|
## (5) the app page's sentence to the household, both languages
|
|
uptime-kuma [hu] http=200
|
|
[Naprak] ...tems:center;gap:.5rem"> <span class="stack-state-badge state-run">Fut</span> <span class="tag tag-ok" title="Ez az alkalmazás a legfrissebb elérhető változatot futtatja.">Naprakész</span> <a href="https://status.enkisfelhom.hu" target="_blank" class="btn btn-sm btn-outline">Megnyitás ↗</a> <a href="/stacks/uptime-kuma/logs" class="btn btn-sm btn
|
|
NEGATIVE CONTROL 'ZZZNOTPRESENT' found: False
|
|
uptime-kuma [en] http=200
|
|
[Up to date] ...align-items:center;gap:.5rem"> <span class="stack-state-badge state-run">Running</span> <span class="tag tag-ok" title="This app is running the newest version available.">Up to date</span> <a href="https://status.enkisfelhom.hu" target="_blank" class="btn btn-sm btn-outline">Open ↗</a> <a href="/stacks/uptime-kuma/logs" class="btn btn-sm btn-out
|
|
NEGATIVE CONTROL 'ZZZNOTPRESENT' found: False
|
|
|
|
## (6) done or HELD, and did a journal entry survive?
|
|
ended: update_phase=done, update_phase_label='Frissitve', updating=false, hold_reason=none
|
|
controller log: 'update uptime-kuma: DONE in 1m0s'
|
|
POSITIVE CONTROL: catalog-cache present -> this IS the controller data dir
|
|
ls: cannot access '/var/lib/docker/volumes/felhom-controller-data/_data/data/update-journal.json': No such file or directory
|
|
-> no journal entry survived the reboot.
|
|
|
|
## STOP CONDITIONS — none tripped
|
|
seeded data gone/unreadable : NO (monitor read back by name)
|
|
updating:true that never clears : NO (cleared 5 s after boot)
|
|
journal surviving a 2nd reboot : N/A, no journal survived the 1st
|
|
pin naming one version, container another : NO
|
|
resumed update retrying in a loop : NO (one resume, one success)
|
|
HOLD with no household sentence : N/A, no hold
|
|
|
|
## SEPARATE DEFECT FOUND WHILE SEEDING (not an update-arc finding, filed here so it is not lost)
|
|
uptime-kuma 2.4.0 first boot sits in its SETUP-DATABASE wizard:
|
|
[SETUP-DATABASE] INFO: Starting Setup Database
|
|
[SETUP-DATABASE] INFO: Waiting for user action...
|
|
The main socket.io server never starts until a database type is chosen. The Felhom controller
|
|
nevertheless reported the app RUNNING and HEALTHY — its probe is `http :3001` and the wizard
|
|
answers 302 on that port. So the box tells the household the monitoring app is fine while it is
|
|
actually parked on an un-passed wizard, with no monitors and no login.
|
|
Passed here through the app's own front door: POST /setup-database {"dbConfig":{"type":"sqlite"}}
|
|
-> {"ok":true}. A catalog fix would pin the DB type at deploy time so first boot never stops.
|