Files
felhom.eu/documentation/audits/update-arc-gaps-2026-09-21/07-scenarioC-controller-only-restart.txt
T
admin c85262111c
gates / gates (push) Successful in 23s
The update arc's two missing measurements, the lock, and the floor to 0.260.0
Part 0 — floor raised to 0.260.0, MinAgent 0.131.0 declared. 3 boxes below, all
down or blocked; both demo boxes SERVED.

Part 1 (R-610) — the DANGEROUS power cut, measured three times with three apps and
two cut mechanisms. All ended honest: resumed, completed, and pinned/installed/live
compose/docker inspect all agreed. vikunja's 2.6.0 migration had ALREADY run 0.64 s
after the cut decision and the seeded data read back intact — so the branch that is
one step from old-binary-on-migrated-database is now evidence, not argument.
Instrument limit stated: `starting` lasts under a second; all three landed in
`verifying`, which RecoverUpdates handles in the same branch.

Part 3 (R-611) — the night the previous session skipped without saying so. An app
updated with nobody pressing anything; a terminally-refused app was pressed exactly
once and never again over three passes. The unattended HOLD was NOT produced: the
within-a-major rule correctly refused the broken edge before it was attempted, so
Q4 still rests on the attended hold from slice 4. Said plainly rather than implied.

Rows: closed R-608/609/610/611; opened R-612 (P1 wishlist unusable on a fresh
install, and its error is a lie), R-613 (uptime-kuma healthy on its setup wizard),
R-614 (stale update phase survives a redeploy). R-520's pointer corrected.

Catalog: two drill pairs, both reverted; every image line byte-identical to
ff9717d3. The alpine:3.20 negative control a security review flagged is cleared.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 15:00:09 +02:00

77 lines
5.2 KiB
Plaintext

# 07 — SCENARIO C: the CONTROLLER-ONLY restart during an app update (wishlist v0.66.0 -> v0.67.0)
#
# Run by the main session on 2026-09-21. This is the scenario that matters most for R-608: it is
# EXACTLY what a controller self-update does to a running app update — the controller container
# restarts, the app's own containers keep running. Guest 9202, controller v0.260.0 (the lock is NOT
# in this build; that is deliberate — this measures the behaviour the lock is meant to make moot).
== TIMESTAMP TABLE ==
poll cadence : 200 ms
app confirmed healthy : 12:35:53.969 UTC (state=running health=True)
Update pressed : 12:35:53.971 UTC -> 202 "Frissites elindult"
phase AT THE DECISION : "starting" observed 12:36:33.136 UTC
cut command : ssh demo-hp 'pct exec 9202 -- systemctl restart felhom-controller-bootstrap.service'
cut command LATENCY : 1 675 ms (rc=0)
phase the box RECORDED : "verifying" (from the controller's own recovery line)
-> THE SAME INSTRUMENT FINDING AS A AND B, REPRODUCED WITH A DIFFERENT AND MUCH FASTER CUT.
The decision was taken at `starting` and the box still recorded `verifying`. A controller
restart returns in 1.7 s where `pct stop` took 3.0-3.8 s, and `starting` STILL could not be
caught. Measured on this box, `starting` lasts well under a second for these apps.
**`RecoverUpdates` handles `starting` and `verifying` in ONE branch (`update.go:909`), so all
three scenarios exercise the same recovery arm** — the arm that says "the new version may have
run". That is the arm the brief wanted measured, and it is measured three times.
== THE APP'S OWN CONTAINERS DURING THE CONTROLLER RESTART ==
felhom-controller 0.260.0 Up 1 second (health: starting) <- restarted
wishlist ghcr.io/cmintey/wishlist:v0.67.0 Up 1 second (health: starting) <- the update's own `up -d`
uptime-kuma louislam/uptime-kuma:2.5.0 Up 3 minutes (healthy) <- UNDISTURBED
vikunja vikunja/vikunja:2.6.0 Up 3 minutes <- UNDISTURBED
glance glanceapp/glance:v0.8.5 Up 3 minutes (healthy) <- UNDISTURBED
filebrowser gtstef/filebrowser:1.3.3-stable Up 3 minutes (healthy) <- UNDISTURBED
-> The controller restart does NOT restart the apps. Only the app the update itself was recreating
shows a new uptime, and that is the update's `up -d`, not the restart.
== THE RECOVERY, verbatim ==
2026/09/21 12:36:34 update.go:909: [WARN] [stacks] update recovery: wishlist was interrupted in verifying (started 2026-09-21T12:35:54Z) — the new version may have run; marking it Updating and RESUMING the health wait
2026/09/21 12:36:34 update.go:951: [INFO] [stacks] update wishlist: resuming after a controller restart — `up -d` then the health wait
2026/09/21 12:36:35 update.go:858: [INFO] [stacks] update wishlist: phase verifying
2026/09/21 12:36:45 update.go:652: [INFO] [stacks] update wishlist: healthy after 10s (the app's health check passed)
2026/09/21 12:36:45 update.go:658: [INFO] [stacks] update wishlist: DONE in 51s
== THE FOUR VERSION OBSERVABLES, SIDE BY SIDE ==
pinned_images : {"wishlist": "ghcr.io/cmintey/wishlist:v0.67.0"}
installed_images : {"wishlist": "ghcr.io/cmintey/wishlist:v0.67.0"}
live compose line: image: ghcr.io/cmintey/wishlist:v0.67.0
docker inspect : ghcr.io/cmintey/wishlist:v0.67.0 | running
-> ALL FOUR AGREE.
== END STATE, AND WHAT THE HOUSEHOLD SEES ==
state=running updating=False update_phase=done label='Frissitve'
update_error=(none) hold_reason=(none) health=True
update-journal.json : ABSENT (cleared)
app page, Hungarian : <span class="tag tag-ok" title="Ez az alkalmazas a legfrissebb elerheto valtozatot futtatja.">Naprakesz</span>
app page, English : <span class="tag tag-ok" title="This app is running the newest version available.">Up to date</span>
(accents transliterated here only to keep this file ASCII-searchable; intact in the page)
no `data-update-error` block on either page — there is nothing to apologise for.
SEARCH CONTROLS: POSITIVE 'Naprak' in hu = 1 · POSITIVE 'Up to date' in en = 1 ·
NEGATIVE 'ZZZ-not-present' = 0
== STOP CONDITIONS — none tripped ==
updating stuck true : NO journal surviving : NO pin vs running image : AGREE
retry loop : NO hold with no sentence : N/A (no hold)
data : wishlist's seeded list item was NOT re-read by this session — see the gap note below.
== GAP, STATED ==
The seeded-data read-back through wishlist's front door was not repeated after this scenario. The
measuring agent seeded and read it back BEFORE the drill (02-seeding.txt) and the app is healthy
and serving on v0.67.0 after it, but "the data is still there" is NOT asserted here for C. A and B
both carry a real post-cut read-back; C does not.
== VERDICT ==
A controller-only restart during an app update ends HONEST. The update resumes, completes, the
other apps are untouched, and the household is shown a clean „Naprakesz" with no error. This is
the behaviour v0.261.0's lock makes unnecessary rather than fixes — worth knowing, because it
means the lock is defence in depth, not a repair of something broken.