The update arc's two missing measurements, the lock, and the floor to 0.260.0
gates / gates (push) Successful in 23s

Part 0 — floor raised to 0.260.0, MinAgent 0.131.0 declared. 3 boxes below, all
down or blocked; both demo boxes SERVED.

Part 1 (R-610) — the DANGEROUS power cut, measured three times with three apps and
two cut mechanisms. All ended honest: resumed, completed, and pinned/installed/live
compose/docker inspect all agreed. vikunja's 2.6.0 migration had ALREADY run 0.64 s
after the cut decision and the seeded data read back intact — so the branch that is
one step from old-binary-on-migrated-database is now evidence, not argument.
Instrument limit stated: `starting` lasts under a second; all three landed in
`verifying`, which RecoverUpdates handles in the same branch.

Part 3 (R-611) — the night the previous session skipped without saying so. An app
updated with nobody pressing anything; a terminally-refused app was pressed exactly
once and never again over three passes. The unattended HOLD was NOT produced: the
within-a-major rule correctly refused the broken edge before it was attempted, so
Q4 still rests on the attended hold from slice 4. Said plainly rather than implied.

Rows: closed R-608/609/610/611; opened R-612 (P1 wishlist unusable on a fresh
install, and its error is a lie), R-613 (uptime-kuma healthy on its setup wizard),
R-614 (stale update phase survives a redeploy). R-520's pointer corrected.

Catalog: two drill pairs, both reverted; every image line byte-identical to
ff9717d3. The alpine:3.20 negative control a security review flagged is cleared.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-21 15:00:09 +02:00
parent d19f07ea04
commit c85262111c
18 changed files with 1300 additions and 76 deletions
@@ -227,6 +227,20 @@ window that starts at 02:30 and ends at 05:00 sits on top of the freshest copy o
new mechanism. **If nothing is decided:** Slice 6 cannot be built at all — every other question below
is downstream of this one.
**⚠ THE WINDOW CONTAINS 04:30, AND 04:30 IS WHEN THE BOX UPDATES ITSELF.** Found 2026-09-21 (R-608) by
reading the clock rather than by a failure: the controller self-updates daily at
`self_update.auto_update_time`, **default 04:30**, and again from `MaybeAutoUpdate` after ANY hub
report once a floor sits above the box — so **at any hour, not only at 04:30**. Either path restarts
the controller container, which is the supervisor of a running app update.
**What v0.261.0 now guarantees, so this question can be answered without also solving that one:** the
two cannot overlap in either direction. The controller defers its own swap while a guarded app update
is in flight (retrying on the next report, exactly as it already did for a running backup), and
`UpdatePreflight` refuses `self_updating` while a swap is in progress. **The lock does not latch** — a
held app does not block the controller's updates for ever. So the window may contain 04:30; the two
jobs will queue behind one another rather than meet. **What it does NOT do is reorder them**: if the
operator prefers the box to take its own update first, that is a scheduling choice still open here.
### Q2 — May an automatic update run on a bind-data app when no copy holds its FILES?
*The button's rule and the automatic rule can differ. Should they?*
@@ -289,6 +303,22 @@ major gets automated by accident.
`reason: update_failed`), plus one mail. **If nothing is decided:** the safe default is no automatic
update at all, because a hold nobody is told about is worse than a version nobody moved.
**MEASURED 2026-09-21, and the honest answer is that HALF of this is still unmeasured.** The
unattended night ran (`audits/update-arc-gaps-2026-09-21/09-unattended-night.md`). What it proved:
an app updates itself end to end with nobody pressing anything; one app takes **51 s – 1 m 26 s**
including the health wait; the caller needs **no new controller code**, only the existing guarded
Update plus `UpdateRefusal.Reason` on the wire (v0.261.0, R-609); and **a terminally-refused app is
pressed exactly ONCE and never again** — four apps, three passes, proven.
**What it did NOT produce is a HOLD, and the reason is instructive rather than a failure of the
run.** The only failing edge available was `vikunja → alpine:3.20`, and the caller **correctly
refused to attempt it**: different repositories cannot be ordered, so the edge is "across" and
belongs to a human by decision 3. **The rule that makes automatic updates safe is the same rule that
refuses the obvious way to break one.** Measuring the unattended hold needs an edge that PASSES the
within-a-major test and still fails its health check — same repository, same major, a tag that
starts and does not serve — which probably means a purpose-built image rather than a catalog move.
**So this question still rests on the ATTENDED hold measured in slice 4 (v0.238.0, Scenario F).**
### Q5 — PostgreSQL: what has to exist before the catalog may move `postgres:16` to `17`?
*Eleven templates, and the image performs no conversion — it refuses to start on an older major's
@@ -682,7 +712,23 @@ window and the per-app switch, and calling `Manager.StartGuardedUpdate`. **It mu
update's health wait is the defect that release fixed, and a second unattended caller is exactly the
shape that finds it again.
**Ships behind `auto_update: off` with no UI until the operator answers Q1.**
**The in-process caller reads `UpdateRefusal.Reason`, and the split is measured, not assumed**
(v0.261.0, R-609): `busy`, `updating`, `deploying`, `migrating` and `self_updating` are **transient**
— try again on the next pass; `held` and `downgrade` are **terminal** — never press that app again
until a person acts; `memory`, `disk` and `no_backup` need a person and should be surfaced, not
retried. **Before the reason reached the wire the only safe readings were "give up on everything" or
"press for ever"**, which is why this is listed as a dependency of the slice rather than a detail of
it. A working caller in this exact shape exists as evidence, not product:
`audits/update-arc-gaps-2026-09-21/unattended-caller.py`.
**Ships behind `app_update.unattended: false` with no UI until the operator answers Q1.**
⚠ **NOT `auto_update` — that name is TAKEN, and by the very thing this must not collide with.**
`self_update.auto_update` / `self_update.auto_update_time` (`config/config.go` L280-281, default
**04:30** at L422) are the CONTROLLER's own update. An earlier draft of this section said Slice 6
"ships behind `auto_update: off`"; two settings with that name, one meaning the controller and one
meaning apps, is the kind of collision that is only discovered by an operator who turned off the
wrong one. The app-scoped key is `app_update.*`.
### 6.3 Slice 7, as it would be built (OPEN — R-451; needs Q7)
@@ -715,10 +761,10 @@ headlessly (R-460).
| leg | what | cost |
|---|---|---|
| A | the **15 database services** — 4 MariaDB + 11 PostgreSQL, across 14 apps by the substring rule plus `adventurelog`'s postgis — one edge each, fixture per app | **15–25 CC-hours**, dominated by seed routes; ~30 min machine time at the median; ~25 GB |
| B | one **power cut mid-update** on a real version change, in `pulling` and again in `starting` | 1–2 CC-hours (R-520 — the first half is measured in this session) |
| B | ~~one power cut mid-update~~ **DONE 2026-09-21 (R-610)** — measured THREE times, two cut mechanisms, three apps: `pulling` (R-520) and the dangerous post-start case three times over. All ended honest; vikunja's 2.6.0 migration had already run when the power went and the data read back intact. **What remains: a cut landing inside `starting` itself** (it lasts well under a second; needs an in-process fault injector, not a faster shell) | 0 — spent |
| C | one **PostgreSQL `pg_upgrade` rehearsal**, the Q5 edge, on one app before any of the eleven | 3–4 CC-hours |
| D | one **downgrade refusal** | **already done** — v0.260.0, proven live 2026-09-21 |
| E | the **automatic night** on a throwaway: one app, one real catalog step, inside a simulated window, with the guard and the hold; then the same edge made to fail → HOLD, the event, no retry loop | 2–3 CC-hours |
| E | ~~the automatic night~~ **MOSTLY DONE 2026-09-21 (R-611)** — the success night and the no-retry proof both measured. **What remains: the unattended HOLD**, which needs an edge that passes the within-a-major test and still fails health (see Q4) | ~1 CC-hour + a purpose-built image |
| F | the remaining **38 apps**, through the nightly rotation as decision 6 directs | ~1 app/night; fixtures amortised |
**Total for legs A–E: roughly 21–34 CC-hours**, plus ~25–30 GB of images on a scratch host. Legs C
@@ -0,0 +1,112 @@
# Driving the guest-9202 controller API from DooPlex — the working recipe (2026-09-21)
Controller v0.260.0 in LXC guest **9202 `demo-hp-scratch`** on host `demo-hp`.
## The one surprise that saves the most time
**Guest 9202 is directly reachable from DooPlex on the home LAN at `192.168.0.114`.**
You do NOT have to `ssh demo-hp` + `pct exec` to drive the API — that was only needed for
`docker`/`pct`. Curling straight from DooPlex is what makes a 200 ms poll loop possible at all.
Still mandatory: the **`Host: felhom.enkisfelhom.hu`** header (without it every route 404s with the
public page) and **`-k`** (self-signed cert on traefik). Port 443. `127.0.0.1:8080` inside the guest
is NOT open — traefik is the only door.
## 0. The password, file→file, never through stdout
```bash
python3 /mnt/5_hdd/felhom.eu/git/felhom.eu/scripts/read_credential.py PASSWORD /tmp/.ctlpw
chmod 600 /tmp/.ctlpw # expect 13 bytes; 15 means it is wearing its quotes
```
## 1. Log in and scrape the session CSRF (both are needed for any POST)
curl's cookie jar drops `felhom_session`, so dump the headers and grep it out by hand.
The CSRF meta tag's closing quote must be stripped or you get a 65-char token that silently
mismatches — check the length is exactly 64.
```bash
S=/tmp/ctl # any scratch dir
mkdir -p $S
B="https://192.168.0.114"
H='Host: felhom.enkisfelhom.hu'
PW=$(cat /tmp/.ctlpw)
curl -sk -D $S/hdr.txt -o /dev/null -H "$H" -X POST --data-urlencode "password=$PW" "$B/login"
head -1 $S/hdr.txt # expect: HTTP/2 302
grep -oiE 'felhom_session=[A-Za-z0-9._-]+' $S/hdr.txt | head -1 > $S/sess.txt # ~79 chars
curl -sk -L -H "$H" -H "Cookie: $(cat $S/sess.txt)" "$B/" -o $S/home.html
grep -oE '<meta name="csrf-token" content="[^"]+"' $S/home.html | head -1 \
| sed 's/.*content="//;s/"$//' > $S/csrf.txt
echo "csrf len: $(wc -c < $S/csrf.txt)" # expect 65 = 64 + newline
```
## 2. The call helper — `c.sh GET /api/stacks` / `c.sh POST /api/sync '{...}'`
```bash
cat > $S/c.sh <<'EOF'
#!/bin/bash
S=/tmp/ctl
B="https://192.168.0.114"
H='Host: felhom.enkisfelhom.hu'
M=$1; P=$2; D=$3
SESS=$(cat $S/sess.txt); CT=$(cat $S/csrf.txt)
if [ "$M" = GET ]; then
curl -sk -H "$H" -H "Cookie: $SESS" "$B$P"
else
curl -sk -H "$H" -H "Cookie: $SESS" -H "X-CSRF-Token: $CT" \
-H "Content-Type: application/json" -X "$M" ${D:+--data "$D"} "$B$P"
fi
EOF
chmod +x $S/c.sh
```
Worked endpoints (all verified today):
| call | meaning |
|---|---|
| `c.sh GET /api/stacks` | every stack; `app_config.pinned_images`, `app_config.installed_images`, `template_images`, `catalog_images`, `updating`, `update_phase` |
| `c.sh GET /api/stacks/vikunja` | one stack, same shape — this is the 200 ms poll target |
| `c.sh POST /api/stacks/<n>/deploy '{"values":{"DOMAIN":"enkisfelhom.hu","SUBDOMAIN":"tasks"}}'` | real deploy path, answers 202 |
| `c.sh POST /api/stacks/<n>/update` | the Update button |
| `c.sh POST /api/sync` | pull the catalog git clone |
| `c.sh POST /api/stacks/rescan` | **always run this after a sync before reading any badge** (R-607) |
| `c.sh GET /api/stacks/<n>/logs?lines=200` | app container log |
`pinned_images` and `installed_images` live under **`app_config`**, not at the top level — that
cost a few minutes.
## 3. Reading the customer's app page, both languages
```bash
curl -sk -H "$H" -H "Cookie: $(cat $S/sess.txt)" "$B/app/vikunja" # Hungarian
curl -sk -H "$H" -H "Cookie: $(cat $S/sess.txt)" "$B/app/vikunja?lang=en" # English
```
## 4. Shell into the guest (for docker / pct only)
```bash
ssh -o StrictHostKeyChecking=accept-new demo-hp "pct exec 9202 -- bash -c '<cmd>'" 2>/dev/null
```
`2>/dev/null` drops the perl locale warnings. For anything with awkward quoting, pipe a script:
```bash
cat <<'EOF' | ssh -o StrictHostKeyChecking=accept-new demo-hp \
'cat > /tmp/c.sh; pct push 9202 /tmp/c.sh /tmp/c.sh >/dev/null 2>&1; pct exec 9202 -- bash /tmp/c.sh' 2>/dev/null
docker ps --format '{{.Names}}\t{{.Image}}'
EOF
```
## 5. Paths inside the running guest
- controller data dir (journal, catalog cache):
`/var/lib/docker/volumes/felhom-controller-data/_data/data/`
- `update-journal.json` — present only while an update is in flight
- `catalog-cache/` — a git clone; `git -C … log --oneline -1` tells you what the box actually has
- stacks: `/opt/docker/stacks/<app>/docker-compose.yml` (the LIVE rendered file) and `app.yaml`
## 6. App front doors on 9202 (Host header per app, same IP)
`tasks.` vikunja · `status.` uptime-kuma · `wishes.` wishlist · `dashboard.` glance — all
`…enkisfelhom.hu` against `https://192.168.0.114`.
@@ -0,0 +1,45 @@
# 01 — baseline after deploying the four apps at today's pins
# guest 9202 (demo-hp-scratch), controller 0.260.0, 2026-09-21T12:14:38Z
== read at 2026-09-21T12:14:38.510303Z
glance state=running deployed=True deploying=False updating=False phase=- label=-
update_error=-
hold_reason =-
pinned ={"glance": "glanceapp/glance:v0.8.5"}
installed={"glance": {"ref": "glanceapp/glance:v0.8.5", "digest": "sha256:32ab73d80f2b8b5fb0735b0431deb36b93fbb6b2fb43592449b0178c8b83e350", "at": "2026-09-21T12:11:56Z"}}
template ={"glance": "glanceapp/glance:v0.8.5"}
catalog ={"glance": "glanceapp/glance:v0.8.5"}
health =healthy=True
uptime-kuma state=running deployed=True deploying=False updating=False phase=done label=Frissítve
update_error=-
hold_reason =-
pinned ={"uptime-kuma": "louislam/uptime-kuma:2.4.0"}
installed={"uptime-kuma": {"ref": "louislam/uptime-kuma:2.4.0", "digest": "sha256:91e963bfda569ba115206e843febb446f473ab525add4e08b2b9e3beffa16985", "at": "2026-09-21T12:11:45Z"}}
template ={"uptime-kuma": "louislam/uptime-kuma:2.4.0"}
catalog ={"uptime-kuma": "louislam/uptime-kuma:2.4.0"}
health =healthy=True
vikunja state=running deployed=True deploying=False updating=False phase=- label=-
update_error=-
hold_reason =-
pinned ={"vikunja": "vikunja/vikunja:2.3.0"}
installed={"vikunja": {"ref": "vikunja/vikunja:2.3.0", "digest": "sha256:f6b80393c1998cd5cd0dc38d24762c59ab4c10000a6f1032ef5b554e262cab93", "at": "2026-09-21T12:11:44Z"}}
template ={"vikunja": "vikunja/vikunja:2.3.0"}
catalog ={"vikunja": "vikunja/vikunja:2.3.0"}
health =healthy=True
wishlist state=running deployed=True deploying=False updating=False phase=- label=-
update_error=-
hold_reason =-
pinned ={"wishlist": "ghcr.io/cmintey/wishlist:v0.66.0"}
installed={"wishlist": {"ref": "ghcr.io/cmintey/wishlist:v0.66.0", "digest": "sha256:073ab4de0f27a93a79410172bedfa3947bed4718f050396547516def790f1f49", "at": "2026-09-21T12:12:17Z"}}
template ={"wishlist": "ghcr.io/cmintey/wishlist:v0.66.0"}
catalog ={"wishlist": "ghcr.io/cmintey/wishlist:v0.66.0"}
health =healthy=True
## docker ps
wishlist ghcr.io/cmintey/wishlist:v0.66.0 Up 2 minutes (healthy)
glance glanceapp/glance:v0.8.5 Up 2 minutes (healthy)
uptime-kuma louislam/uptime-kuma:2.4.0 Up 2 minutes (healthy)
vikunja vikunja/vikunja:2.3.0 Up 2 minutes
felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.260.0 Up About an hour (healthy)
filebrowser gtstef/filebrowser:1.3.3-stable Up About an hour (healthy)
traefik traefik:v3.6.7 Up About an hour
@@ -0,0 +1,37 @@
# 02 — seeding each app through its own front door, BEFORE any bump
# guest 9202, 2026-09-21T12:24:42Z
## vikunja — seeded via its REST API (/api/v1/register, /login, PUT /projects, PUT /projects/2/tasks)
vikunja version: v2.3.0
TASK 1 'SEED-VIKUNJA-CANARY-9f3c1e-20260921' desc= 'power-cut drill canary'
POSITIVE CONTROL: seed present = True
NEGATIVE CONTROL: absent string present = False
## uptime-kuma — seeded via its own socket.io front door (the same API the browser uses)
# NOTE: uptime-kuma 2.4.0 first boot sits in the SETUP-DATABASE wizard ('Waiting for user action').
# The controller reported the app running + healthy anyway. The wizard was passed through its own
# front door: POST /setup-database {"dbConfig":{"type":"sqlite"}} -> {"ok":true}.
MONITORS: ["SEED-KUMA-CANARY-4a91c7-20260921"]
POSITIVE CONTROL: seed present = true
NEGATIVE CONTROL: absent name present = false
read-back finished
## wishlist — seeded via POST /signup then POST /lists/<id>/create-item
# NOTE: on first boot the image's own 'pnpm prisma db seed' was OOM-Killed at mem_limit 128M,
# so the Role/Group rows were missing and EVERY signup failed with a misleading
# 'User with username or email already exists' (real cause: FOREIGN KEY constraint).
# Repaired by running the image's own seed once (docker update --memory 512m, run seed, back to 128m).
wishlist version banner: 308
list page code=200
POSITIVE CONTROL occurrences of seed name: 2
NEGATIVE CONTROL occurrences of absent name: 0
context: '-[--><!--[-1--><div data-scope="dialog" data-part="title" id="dialog:s11:title" class="truncate text-xl font-bold text-wrap wrap-break-word md:text-2xl"><!---->SEED-WISHLIST-CANARY-7b2d44-20260921<!----></div><!--]--><!--]--> <!--[--><!--[-1--><button data-scope="dialog" data-par'
## glance — no user data
glance has NO user data (catalog app_info declares none; the only persistent state is the
first-boot seeded /app/config/glance.yml). Observable recorded instead:
d9104faa1890d25cd77ed62eb2271da5 /app/config/glance.yml
1
glance HTTP 200 size=7842
POSITIVE CONTROL (ASCII fragment "Kezd" on the page): 4
NEGATIVE CONTROL: 0
@@ -0,0 +1,105 @@
# 03 — DRILL commit pushed, catalog synced, all four read 'update available'
# 2026-09-21T12:27:11Z
## catalog commit
573e41f DRILL: move four app pins for the power-cut update-arc measurement
573e41f DRILL: move four app pins for the power-cut update-arc measurement
templates/glance/.felhom.yml | 2 +-
templates/glance/docker-compose.yml | 2 +-
templates/uptime-kuma/docker-compose.yml | 2 +-
templates/vikunja/.felhom.yml | 2 +-
templates/vikunja/docker-compose.yml | 2 +-
templates/wishlist/.felhom.yml | 2 +-
templates/wishlist/docker-compose.yml | 2 +-
7 files changed, 7 insertions(+), 7 deletions(-)
## the box's catalog clone after POST /api/sync
573e41f DRILL: move four app pins for the power-cut update-arc measurement
## POST /api/sync said: Sablonok frissitve - frissitve: glance, vikunja, wishlist
## (uptime-kuma NOT listed - its .felhom.yml was already at catalog_since 2026-09-21 from the
## earlier session, and the syncer names only files whose content it rewrote. R-607 trap:
## a POST /api/stacks/rescan was run straight after, and only then were the badges read.)
## POST /api/stacks/rescan said: Rescan completed: 55 stacks found
## the four observables per app, AFTER sync+rescan
== read at 2026-09-21T12:27:12.574916Z
glance state=running deployed=True deploying=False updating=False phase=- label=-
update_error=-
hold_reason =-
pinned ={"glance": "glanceapp/glance:v0.8.5"}
installed={"glance": {"ref": "glanceapp/glance:v0.8.5", "digest": "sha256:32ab73d80f2b8b5fb0735b0431deb36b93fbb6b2fb43592449b0178c8b83e350", "at": "2026-09-21T12:11:56Z"}}
template ={"glance": "glanceapp/glance:v0.8.5"}
catalog ={"glance": "glanceapp/glance:v0.8.6"}
health =healthy=True
uptime-kuma state=running deployed=True deploying=False updating=False phase=done label=Frissítve
update_error=-
hold_reason =-
pinned ={"uptime-kuma": "louislam/uptime-kuma:2.4.0"}
installed={"uptime-kuma": {"ref": "louislam/uptime-kuma:2.4.0", "digest": "sha256:91e963bfda569ba115206e843febb446f473ab525add4e08b2b9e3beffa16985", "at": "2026-09-21T12:11:45Z"}}
template ={"uptime-kuma": "louislam/uptime-kuma:2.4.0"}
catalog ={"uptime-kuma": "louislam/uptime-kuma:2.5.0"}
health =healthy=True
vikunja state=running deployed=True deploying=False updating=False phase=- label=-
update_error=-
hold_reason =-
pinned ={"vikunja": "vikunja/vikunja:2.3.0"}
installed={"vikunja": {"ref": "vikunja/vikunja:2.3.0", "digest": "sha256:f6b80393c1998cd5cd0dc38d24762c59ab4c10000a6f1032ef5b554e262cab93", "at": "2026-09-21T12:11:44Z"}}
template ={"vikunja": "vikunja/vikunja:2.3.0"}
catalog ={"vikunja": "vikunja/vikunja:2.6.0"}
health =healthy=True
wishlist state=running deployed=True deploying=False updating=False phase=- label=-
update_error=-
hold_reason =-
pinned ={"wishlist": "ghcr.io/cmintey/wishlist:v0.66.0"}
installed={"wishlist": {"ref": "ghcr.io/cmintey/wishlist:v0.66.0", "digest": "sha256:073ab4de0f27a93a79410172bedfa3947bed4718f050396547516def790f1f49", "at": "2026-09-21T12:12:17Z"}}
template ={"wishlist": "ghcr.io/cmintey/wishlist:v0.66.0"}
catalog ={"wishlist": "ghcr.io/cmintey/wishlist:v0.67.0"}
health =healthy=True
## the badge, both languages — ASCII-only fragments, with a negative control
vikunja[hu] http=200 size=43698 | fragments hit: ['Friss'] | NEGATIVE CONTROL ZZZ-NOT-IN-ANY-PAGE-ZZZ hit: False
BADGE: 'Fut'
BADGE: '~50M RAM'
BADGE: 'productivity'
BADGE: 'Pi kompatibilis'
vikunja[en] http=200 size=42909 | fragments hit: ['Update available'] | NEGATIVE CONTROL ZZZ-NOT-IN-ANY-PAGE-ZZZ hit: False
BADGE: 'Running'
BADGE: '~50M RAM'
BADGE: 'productivity'
BADGE: 'Runs on Pi'
uptime-kuma[hu] http=200 size=43881 | fragments hit: ['Friss'] | NEGATIVE CONTROL ZZZ-NOT-IN-ANY-PAGE-ZZZ hit: False
BADGE: 'Fut'
BADGE: '~50M RAM'
BADGE: 'dashboard'
BADGE: 'Pi kompatibilis'
uptime-kuma[en] http=200 size=43092 | fragments hit: ['Update available'] | NEGATIVE CONTROL ZZZ-NOT-IN-ANY-PAGE-ZZZ hit: False
BADGE: 'Running'
BADGE: '~50M RAM'
BADGE: 'dashboard'
BADGE: 'Runs on Pi'
wishlist[hu] http=200 size=43666 | fragments hit: ['Friss'] | NEGATIVE CONTROL ZZZ-NOT-IN-ANY-PAGE-ZZZ hit: False
BADGE: 'Fut'
BADGE: '~30M RAM'
BADGE: 'home'
BADGE: 'Pi kompatibilis'
wishlist[en] http=200 size=42878 | fragments hit: ['Update available'] | NEGATIVE CONTROL ZZZ-NOT-IN-ANY-PAGE-ZZZ hit: False
BADGE: 'Running'
BADGE: '~30M RAM'
BADGE: 'home'
BADGE: 'Runs on Pi'
glance[hu] http=200 size=43734 | fragments hit: ['Friss'] | NEGATIVE CONTROL ZZZ-NOT-IN-ANY-PAGE-ZZZ hit: False
BADGE: 'Fut'
BADGE: '~20M RAM'
BADGE: 'dashboard'
BADGE: 'Pi kompatibilis'
glance[en] http=200 size=42966 | fragments hit: ['Update available'] | NEGATIVE CONTROL ZZZ-NOT-IN-ANY-PAGE-ZZZ hit: False
BADGE: 'Running'
BADGE: '~20M RAM'
BADGE: 'dashboard'
BADGE: 'Runs on Pi'
## the exact sentence, verbatim (vikunja)
HU: <span class="tag tag-warn" title="Ujabb valtozat erheto el ehhez az alkalmazashoz. A frissites inditasahoz nyomd meg a Frissites gombot.">Frissites elerheto — ma</span>
(accents intact in the page; transliterated here only to keep this file ASCII-searchable)
EN: <span class="tag tag-warn" title="A newer version of this app is available. Select the Update button to start it.">Update available — today</span>
@@ -0,0 +1,29 @@
# 04b — INDEPENDENT CORROBORATION of scenario A, captured by the COORDINATOR
#
# WHY THIS FILE EXISTS, stated plainly: 25 minutes passed with no evidence written while the box
# already showed the update complete, so the main session took its own read-only capture rather than
# risk losing the measurement (R-320 — evidence off before it can be lost). The measuring agent's own
# file `04-scenarioA-starting-cut.txt` then landed and is FULLER and BETTER than this one: it has the
# timestamp table, the journal read off the stopped guest, and the seeded-data read-back.
#
# THIS FILE IS NOT A SECOND MEASUREMENT. It is the same box, read independently ~90 minutes later by
# a different reader with a different session. Its only value is that it agrees.
READ AT: 2026-09-21, after the scenario completed. Source: live guest 9202, read-only.
AGREES WITH 04-scenarioA-starting-cut.txt ON:
* the recovery line — interrupted in `verifying`, "the new version may have run; marking it
Updating and RESUMING the health wait", resumed, `up -d`, healthy after 0s, DONE in 1m26s;
* all FOUR version observables reading vikunja/vikunja:2.6.0 —
pinned_images, installed_images, the live compose `image:` line, and `docker inspect`
(created 2026-09-21T12:28:26Z, state running);
* end state: state=running, updating=False, update_phase=done, label 'Frissitve',
no update_error, no hold_reason, health_probe.healthy=True;
* the update journal ABSENT (cleared), with a positive control that the directory searched is the
real one — catalog-cache, debug-ring.log, encryption.key, metrics.db*, settings.json* beside it.
WHAT THIS FILE COULD NOT DO, and why — because a gap recorded is worth more than a gap hidden:
the seeded-data read-back. 02-seeding.txt records the canary task name but not the account it was
created under, and the coordinator would not guess credentials or read vikunja's database behind
the app's back. "Through the front door" is the condition the measurement is worth anything under.
The measuring agent had the credentials and did it: the task read back unchanged.
@@ -0,0 +1,129 @@
# 04 — SCENARIO A: vikunja 2.3.0 -> 2.6.0, cut decided in `starting`
# guest 9202 (demo-hp-scratch), controller 0.260.0, 2026-09-21
#
# VERDICT: the box ended HONEST. The update resumed after the reboot and finished `done`;
# all four version observables agree on 2.6.0; the seeded task read back unchanged; the
# journal is gone; the household sentence is „Naprakesz" / "Up to date".
#
# THE ONE THING THAT DID NOT GO AS THE BRIEF ASSUMED — stated first because it changes
# how capture (1) should be read:
# The DECISION was taken in `starting` (12:28:26.244 UTC, poll cadence ~21 ms).
# But `pct stop 9202` took 3 759 ms to return, and the update moved on inside that window.
# The journal read OFF the stopped guest proves where the box actually died:
# phase "verifying", not "starting".
# `starting` on this box lasts about 0.6 s (12:28:26.244 -> vikunja's own 14:28:26.881+02:00
# migration line). NO instrument available here can land a guest kill inside it: the poll is
# fast enough (21 ms), the CUT COMMAND is not (3.8 s).
# This costs nothing for the measurement, because RecoverUpdates handles `starting` and
# `verifying` in the SAME branch (update.go:909 `case UpdatePhaseStarting, UpdatePhaseVerifying`).
# Scenario A and Scenario B therefore exercise one recovery arm, not two. Recorded, not hidden.
## (1) TIMESTAMP TABLE
poll target : GET /api/stacks/vikunja, HTTPS keep-alive from DooPlex
measured poll cadence : ~21 ms (per-sample HTTP cost 1.1-1.3 ms + 20 ms sleep)
Update pressed : 12:28:23.075 UTC (POST /api/stacks/vikunja/update -> {"accepted":true})
phase AT THE DECISION : "starting" observed 12:28:26.244 UTC
cut command : ssh -S <prewarmed master> demo-hp 'pct stop 9202'
cut command LATENCY : 3 759 ms (returned 12:28:30.006 UTC, rc=0, no output)
phase the box DIED in : "verifying" (from update-journal.json read off the stopped guest)
guest restarted : pct start 9202 returned after 3.6 s; controller up 12:29:48 UTC
total app downtime : ~79 s (vikunja stopped 12:28:26.9 .. restarted 12:29:47.3 UTC)
## (1b) the phase trace, verbatim from the poller
12:28:21.860 updating=False phase='' label='' state=running err=''
12:28:23.088 updating=True phase='safety-dump' label='Adatbázis pillanatkép…' state=running err=''
12:28:23.110 updating=True phase='pinning' label='Új verzió letöltése…' state=running err=''
12:28:23.132 updating=True phase='pulling' label='Új verzió letöltése…' state=running err=''
12:28:26.244 updating=True phase='starting' label='Indítás az új verzióval…' state=running err=''
--- DECISION at 12:28:26.244: phase='starting' -> firing cut: ssh -S /tmp/cm-demohp demo-hp 'pct stop 9202'
--- CUT returned rc=0 after 3759 ms
--- CUT stdout:
--- CUT stderr:
=== TIMESTAMP TABLE (vikunja) ===
poll cadence : measured below
phase at decision : 'starting' at 12:28:26.244 UTC
cut command : ssh -S /tmp/cm-demohp demo-hp 'pct stop 9202'
cut command latency : 3759 ms (returned 12:28:30.006 UTC)
## (1c) update-journal.json read from the STOPPED guest
# The brief's path was WRONG for this guest: /var/lib/lxc/9202/rootfs/var/lib/felhom is an
# EMPTY MOUNTPOINT while the guest is stopped, because mp0 is a separate raw volume
# (nvme-scratch:9202/vm-9202-disk-1.raw). The POSITIVE CONTROL failed there: no catalog-cache/.
/bin/bash: line 131: pct: command not found
# DOES attach mp0, and then the same path is real (catalog-cache/ present).
/bin/bash: line 132: pct: command not found
/bin/bash: line 132: pct: command not found
# Remember to afterwards or refuses.
{
"updates": {
"vikunja": {
"phase": "verifying",
"started_at": "2026-09-21T12:28:23.06699273Z",
"prev_pin": {
"vikunja": "vikunja/vikunja:2.3.0"
},
"prev_compose": "/opt/docker/stacks/vikunja/pre-update-compose.yml",
"prev_applied": "/opt/docker/stacks/vikunja/pre-update-applied.yml",
"proven_copy_at": "2026-09-21T12:26:12Z",
"proven_tier": 1
}
}
}
## (2) RecoverUpdates log lines after the restart, VERBATIM
2026/09/21 12:29:48 update.go:909: [WARN] [stacks] update recovery: vikunja was interrupted in verifying (started 2026-09-21T12:28:23Z) — the new version may have run; marking it Updating and RESUMING the health wait
2026/09/21 12:29:48 update.go:951: [INFO] [stacks] update vikunja: resuming after a controller restart — `up -d` then the health wait
2026/09/21 12:29:48 update.go:858: [INFO] [stacks] update vikunja: phase verifying
2026/09/21 12:29:48 update.go:652: [INFO] [stacks] update vikunja: healthy after 0s (the app's health check passed)
2026/09/21 12:29:49 update.go:658: [INFO] [stacks] update vikunja: DONE in 1m26s
## (3) THE FOUR VERSION OBSERVABLES, SIDE BY SIDE (after recovery)
1_pinned_images : {"vikunja": "vikunja/vikunja:2.6.0"}
2_installed_images : {"vikunja": "vikunja/vikunja:2.6.0"} | digest: {"vikunja": "sha256:417ada6f94e81f0267a"}
(updating=False phase=done label=Frissítve err=- hold=-)
3_live compose line: image: vikunja/vikunja:2.6.0
4_docker inspect : vikunja/vikunja:2.6.0 | RepoDigest: sha256:417ada6f94e81f0267a | started 2026-09-21T12:29:47.287254768Z
-> all four name vikunja/vikunja:2.6.0; installed digest and the running container's
RepoDigest are the same sha256:417ada6f94e81f0267a...
## (4) the seeded data, read back through vikunja's own front door
vikunja version: v2.6.0
TASK 1 'SEED-VIKUNJA-CANARY-9f3c1e-20260921' desc= 'power-cut drill canary'
POSITIVE CONTROL: seed present = True
NEGATIVE CONTROL: absent string present = False
## (4b) the vikunja MIGRATION line, VERBATIM from its container log
3:time=2026-09-21T14:28:26.881+02:00 level=INFO msg="Running migrations…"
5:time=2026-09-21T14:28:26.915+02:00 level=INFO msg="Ran all migrations successfully."
9:time=2026-09-21T14:28:26.922+02:00 level=INFO msg="Vikunja version v2.6.0"
14:time=2026-09-21T14:29:47.713+02:00 level=INFO msg="Running migrations…"
16:time=2026-09-21T14:29:47.734+02:00 level=INFO msg="Ran all migrations successfully."
20:time=2026-09-21T14:29:47.745+02:00 level=INFO msg="Vikunja version v2.6.0"
-> the FIRST migration ran at 14:28:26.881+02:00 = 12:28:26.881 UTC — 0.64 s AFTER the cut
decision and 3.1 s BEFORE pct stop returned. The 2.6.0 schema migration had ALREADY
been applied to the customer's SQLite database when the power went. The pin did not
roll back (that only happens in the pinning/pulling branch), so old binary vs migrated
database never happened here — but it is the failure this branch is one step away from.
## (5) the app page's sentence to the household, both languages
vikunja [hu] http=200
[Naprak] ...tems:center;gap:.5rem"> <span class="stack-state-badge state-run">Fut</span> <span class="tag tag-ok" title="Ez az alkalmazás a legfrissebb elérhető változatot futtatja.">Naprakész</span> <a href="https://tasks.enkisfelhom.hu" target="_blank" class="btn btn-sm btn-outline">Megnyitás ↗</a> <a href="/stacks/vikunja/logs" class="btn btn-sm btn-outl
NEGATIVE CONTROL 'ZZZNOTPRESENT' found: False
vikunja [en] http=200
[Up to date] ...align-items:center;gap:.5rem"> <span class="stack-state-badge state-run">Running</span> <span class="tag tag-ok" title="This app is running the newest version available.">Up to date</span> <a href="https://tasks.enkisfelhom.hu" target="_blank" class="btn btn-sm btn-outline">Open ↗</a> <a href="/stacks/vikunja/logs" class="btn btn-sm btn-outline"
NEGATIVE CONTROL 'ZZZNOTPRESENT' found: False
## (6) done or HELD, and did a journal entry survive?
ended: update_phase=done, update_phase_label='Frissitve', updating=false, hold_reason=none
controller log: 'update vikunja: DONE in 1m26s'
POSITIVE CONTROL: catalog-cache present -> this IS the controller data dir
ls: cannot access '/var/lib/docker/volumes/felhom-controller-data/_data/data/update-journal.json': No such file or directory
-> no journal entry survived the reboot.
## STOP CONDITIONS — none tripped
seeded data gone/unreadable : NO (read back byte-identical)
updating:true that never clears : NO (cleared 1.3 s after boot)
journal surviving a 2nd reboot : N/A, no journal survived the 1st
pin naming one version, container another : NO
resumed update retrying in a loop : NO (one resume, one success)
HOLD with no household sentence : N/A, no hold
@@ -0,0 +1,128 @@
# 05 — SCENARIO B: uptime-kuma 2.4.0 -> 2.5.0, cut during `verifying` (the health wait)
# guest 9202 (demo-hp-scratch), controller 0.260.0, 2026-09-21
#
# VERDICT: the box ended HONEST. The update resumed after the reboot and finished `done`;
# all four version observables agree on 2.5.0; the seeded monitor read back through
# uptime-kuma's own socket.io front door; no journal survived; the household sentence is
# „Naprakesz" / "Up to date".
#
# INSTRUMENT FINDING (this one matters for reading BOTH A and B):
# `pct stop 9202` RETURNS after ~3.0-3.8 s, but the guest stops answering after ~1.1 s.
# Measured here by firing the cut in a THREAD and continuing to poll through it:
# decision 12:32:32.271, last successful API sample 12:32:33.345 (+1 075 ms),
# command return 12:32:35.303 (+3 032 ms).
# So the kill lands early and the rest of the 3 s is teardown. The practical consequence,
# proven in Scenario A: `starting` lasts about 0.6 s on this box, so NO cut driven this way
# can land inside `starting` — by the time the guest dies the update is in `verifying`.
# A and B therefore exercise the SAME RecoverUpdates branch
# (update.go:909 `case UpdatePhaseStarting, UpdatePhaseVerifying`). Stated, not hidden.
## (1) TIMESTAMP TABLE
poll target : GET /api/stacks/uptime-kuma, HTTPS keep-alive from DooPlex
measured poll cadence : ~21 ms (per-sample HTTP cost 1.1-1.3 ms + 20 ms sleep)
Update pressed : 12:32:26.374 UTC (POST /api/stacks/uptime-kuma/update -> accepted)
phase AT THE DECISION : "verifying" observed 12:32:32.271 UTC
cut command : ssh -S <prewarmed master> demo-hp 'pct stop 9202'
cut command LATENCY : 3 016 ms (rc=0, returned 12:32:35.303 UTC, no stdout/stderr)
box stopped ANSWERING : 12:32:33.345 UTC was the LAST successful sample (+1 075 ms)
phase the box DIED in : "verifying" (update-journal.json read off the stopped guest)
guest restarted : pct start 9202; controller up 12:33:21 UTC
update concluded : 12:33:26 UTC, "DONE in 1m0s"
## (1b) the phase trace, verbatim from the poller
12:32:25.147 updating=False phase='' state=running err='' hold=''
12:32:26.373 updating=True phase='checking' state=running err='' hold=''
12:32:26.394 updating=True phase='safety-dump' state=running err='' hold=''
12:32:26.416 updating=True phase='pulling' state=running err='' hold=''
12:32:27.627 updating=True phase='starting' state=running err='' hold=''
12:32:28.748 updating=True phase='starting' state=degraded err='' hold=''
12:32:32.271 updating=True phase='verifying' state=starting err='' hold=''
--- DECISION at 12:32:32.271: phase='verifying' -> firing (in a thread): ssh -S /tmp/cm-demohp demo-hp 'pct stop 9202'
12:32:33.367 ERR ConnectionResetError: [Errno 104] Connection reset by peer
12:32:33.393 ERR TimeoutError: timed out
=== TIMESTAMP TABLE (uptime-kuma) ===
phase at decision : 'verifying' at 12:32:32.271 UTC
cut command : ssh -S /tmp/cm-demohp demo-hp 'pct stop 9202'
cut command latency : 3016 ms (rc=0, returned 12:32:35.303 UTC)
cut stdout/stderr : '' / ''
LAST SUCCESSFUL SAMPLE : 12:32:33.345 UTC <- the box was still answering here
i.e. 1075 ms after the decision
## (1c) update-journal.json read from the STOPPED guest (pct mount 9202 first; pct unmount after)
POSITIVE CONTROL: catalog-cache present -> real dir
--- journal ---
{
"updates": {
"uptime-kuma": {
"phase": "verifying",
"started_at": "2026-09-21T12:32:26.364656602Z",
"prev_pin": {
"uptime-kuma": "louislam/uptime-kuma:2.4.0"
},
"prev_compose": "/opt/docker/stacks/uptime-kuma/pre-update-compose.yml",
"prev_applied": "/opt/docker/stacks/uptime-kuma/pre-update-applied.yml",
"proven_copy_at": "2026-09-21T12:16:12Z",
"proven_tier": 1
}
}
}
## (2) RecoverUpdates log lines after the restart, VERBATIM
2026/09/21 12:33:21 update.go:909: [WARN] [stacks] update recovery: uptime-kuma was interrupted in verifying (started 2026-09-21T12:32:26Z) — the new version may have run; marking it Updating and RESUMING the health wait
2026/09/21 12:33:21 update.go:951: [INFO] [stacks] update uptime-kuma: resuming after a controller restart — `up -d` then the health wait
2026/09/21 12:33:21 update.go:858: [INFO] [stacks] update uptime-kuma: phase verifying
2026/09/21 12:33:26 update.go:652: [INFO] [stacks] update uptime-kuma: healthy after 5s (the app's health check passed)
2026/09/21 12:33:26 update.go:658: [INFO] [stacks] update uptime-kuma: DONE in 1m0s
## (3) THE FOUR VERSION OBSERVABLES, SIDE BY SIDE (after recovery)
1_pinned_images : {"uptime-kuma": "louislam/uptime-kuma:2.5.0"}
2_installed_images : {"uptime-kuma": "louislam/uptime-kuma:2.5.0"} | digest: {"uptime-kuma": "sha256:a8610b3b4c38077922b"}
(updating=False phase=done label=Frissítve err=- hold=-)
3_live compose line: image: louislam/uptime-kuma:2.5.0
4_docker inspect : louislam/uptime-kuma:2.5.0 | RepoDigest: sha256:a8610b3b4c38077922b | started 2026-09-21T12:33:19.417787988Z
-> all four name louislam/uptime-kuma:2.5.0; installed digest and the running
container's RepoDigest are the same sha256:a8610b3b4c38077922b...
## (4) the seeded data, read back through uptime-kuma's OWN front door (socket.io, the same
## API the browser uses), run from inside the container with its own socket.io-client
MONITORS: ["SEED-KUMA-CANARY-4a91c7-20260921"]
POSITIVE CONTROL: seed present = true
NEGATIVE CONTROL: absent name present = false
read-back finished
NOTE: the FIRST read-back attempt, run 12 s after the app came up, TIMED OUT waiting for
the monitorList event — the app was listening but not yet serving. Re-run 20 s later it
answered. Recorded because a single timeout here reads exactly like data loss and is not.
## (5) the app page's sentence to the household, both languages
uptime-kuma [hu] http=200
[Naprak] ...tems:center;gap:.5rem"> <span class="stack-state-badge state-run">Fut</span> <span class="tag tag-ok" title="Ez az alkalmazás a legfrissebb elérhető változatot futtatja.">Naprakész</span> <a href="https://status.enkisfelhom.hu" target="_blank" class="btn btn-sm btn-outline">Megnyitás ↗</a> <a href="/stacks/uptime-kuma/logs" class="btn btn-sm btn
NEGATIVE CONTROL 'ZZZNOTPRESENT' found: False
uptime-kuma [en] http=200
[Up to date] ...align-items:center;gap:.5rem"> <span class="stack-state-badge state-run">Running</span> <span class="tag tag-ok" title="This app is running the newest version available.">Up to date</span> <a href="https://status.enkisfelhom.hu" target="_blank" class="btn btn-sm btn-outline">Open ↗</a> <a href="/stacks/uptime-kuma/logs" class="btn btn-sm btn-out
NEGATIVE CONTROL 'ZZZNOTPRESENT' found: False
## (6) done or HELD, and did a journal entry survive?
ended: update_phase=done, update_phase_label='Frissitve', updating=false, hold_reason=none
controller log: 'update uptime-kuma: DONE in 1m0s'
POSITIVE CONTROL: catalog-cache present -> this IS the controller data dir
ls: cannot access '/var/lib/docker/volumes/felhom-controller-data/_data/data/update-journal.json': No such file or directory
-> no journal entry survived the reboot.
## STOP CONDITIONS — none tripped
seeded data gone/unreadable : NO (monitor read back by name)
updating:true that never clears : NO (cleared 5 s after boot)
journal surviving a 2nd reboot : N/A, no journal survived the 1st
pin naming one version, container another : NO
resumed update retrying in a loop : NO (one resume, one success)
HOLD with no household sentence : N/A, no hold
## SEPARATE DEFECT FOUND WHILE SEEDING (not an update-arc finding, filed here so it is not lost)
uptime-kuma 2.4.0 first boot sits in its SETUP-DATABASE wizard:
[SETUP-DATABASE] INFO: Starting Setup Database
[SETUP-DATABASE] INFO: Waiting for user action...
The main socket.io server never starts until a database type is chosen. The Felhom controller
nevertheless reported the app RUNNING and HEALTHY — its probe is `http :3001` and the wizard
answers 302 on that port. So the box tells the household the monitoring app is fine while it is
actually parked on an un-passed wizard, with no monitors and no login.
Passed here through the app's own front door: POST /setup-database {"dbConfig":{"type":"sqlite"}}
-> {"ok":true}. A catalog fix would pin the DB type at deploy time so first boot never stops.
@@ -0,0 +1,35 @@
# 05b — SCENARIO B (cut in `verifying`, uptime-kuma) — COORDINATOR capture
#
# PROVENANCE: written by the main session, read-only, live from guest 9202, after the measuring
# agent again went a long interval without writing evidence while the box already showed the
# scenario complete. If the agent's own `05-*` file lands it is the fuller record; this exists so
# the measurement could not be lost (R-320). Nothing below is inferred.
== THE RECOVERY, verbatim from the controller log ==
2026/09/21 12:33:21 update.go:909: [WARN] [stacks] update recovery: uptime-kuma was interrupted in verifying (started 2026-09-21T12:32:26Z) — the new version may have run; marking it Updating and RESUMING the health wait
2026/09/21 12:33:21 update.go:951: [INFO] [stacks] update uptime-kuma: resuming after a controller restart — `up -d` then the health wait
2026/09/21 12:33:21 update.go:858: [INFO] [stacks] update uptime-kuma: phase verifying
2026/09/21 12:33:26 update.go:652: [INFO] [stacks] update uptime-kuma: healthy after 5s (the app's health check passed)
2026/09/21 12:33:26 update.go:658: [INFO] [stacks] update uptime-kuma: DONE in 1m0s
== THE FOUR VERSION OBSERVABLES, SIDE BY SIDE ==
pinned_images : {"uptime-kuma": "louislam/uptime-kuma:2.5.0"}
installed_images : {"uptime-kuma": "louislam/uptime-kuma:2.5.0"}
live compose line: image: louislam/uptime-kuma:2.5.0
docker inspect : louislam/uptime-kuma:2.5.0 | running
-> ALL FOUR AGREE. (catalog_images also 2.5.0 — the drill bump was still in place.)
== END STATE ==
state=running updating=False update_phase=done label='Frissitve'
update_error=(none) hold_reason=(none) health_probe.healthy=True
Total time from interruption to done: 1m0s.
== NOT CAPTURED HERE ==
the timestamp table (phase at decision, cut latency), the seeded-monitor read-back through
uptime-kuma's own front door, and the page sentence in both languages. Those need the agent's
poller output and its seeded session. Recorded as GAPS, not as passes.
== VERDICT ON WHAT IS MEASURED ==
The `verifying` cut ends HONEST for a second app, with a different health-check shape (5s rather
than 0s to pass). Resumed, completed, all four observables agree, no hold, no stuck `Updating`.
None of the brief's STOP conditions appeared in what was captured.
@@ -0,0 +1,47 @@
# 06 — THE REVERT IS STILL OWED. Written 2026-09-21 by the measuring session.
## State right now
app-catalog-felhom.eu `main` is at 573e41f59bb3c3f01a04f658f1c02bdfbd427a07
"DRILL: move four app pins for the power-cut update-arc measurement"
The pre-drill commit is ff9717d3794974724e09f8fd58abf058d4cdc2d0
Working tree clean. The drill bump is LIVE on main and on every box that syncs the catalog.
## Why it was not reverted by the session that made it
The operator brief said "Your catalog commits MUST be reverted before you finish."
The coordinating session stood this session down from Scenario C and said it would do the
REVERT itself, because it needs the catalog left bumped until Scenario C and one further
phase are finished. Reverting under a run in progress on the same box would have changed
the catalog beneath that measurement.
So the revert was NOT done here — and this file exists so that is a recorded hand-off and
not a silently dropped fence. If the later phases are finished and no REVERT commit follows
573e41f, THIS IS THE OUTSTANDING ITEM.
## Exactly what the revert must restore (verified, not assumed)
7 files, 7 insertions, 7 deletions — nothing else is in the drill commit:
templates/vikunja/docker-compose.yml vikunja/vikunja:2.6.0 -> 2.3.0
templates/uptime-kuma/docker-compose.yml louislam/uptime-kuma:2.5.0 -> 2.4.0
templates/wishlist/docker-compose.yml ghcr.io/cmintey/wishlist:v0.67.0 -> v0.66.0
templates/glance/docker-compose.yml glanceapp/glance:v0.8.6 -> v0.8.5
templates/vikunja/.felhom.yml catalog_since "2026-09-21" -> "2026-07-18"
templates/glance/.felhom.yml catalog_since "2026-09-21" -> "2026-07-18"
templates/wishlist/.felhom.yml catalog_since "2026-09-21" -> "2026-07-19"
(templates/uptime-kuma/.felhom.yml was ALREADY at catalog_since "2026-09-21" before the drill,
from the earlier session's own drill+revert pair, so it is not in the diff and must NOT be
moved back to an older date.)
The tree object at ff9717d is bef76c8ccb7b2be919e22bc689213ef3199ee86e — a correct revert
must produce a tree identical to that one for these seven paths.
CHECKED (not applied): `git diff 573e41f ff9717d | git apply --check` passes cleanly.
## The command
cd /mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu
git revert --no-edit 573e41f # or apply the reverse diff and commit as "REVERT ..."
python3 scripts/catalog_gates.py --fast # must be all-green BEFORE the push
git push origin main # the pre-push hook re-runs the gates; never --no-verify
# then on any box that must see it back: POST /api/sync AND POST /api/stacks/rescan (R-607)
## A warning for whoever presses Update after the revert
vikunja and uptime-kuma on guest 9202 are now INSTALLED AHEAD of the reverted catalog
(2.6.0 vs 2.3.0, 2.5.0 vs 2.4.0). Controller v0.260.0 refuses that as a downgrade —
"update REFUSED (downgrade): installed is provably NEWER than the catalog on every differing
service" — and the badge reads „Naprakesz". That is CORRECT behaviour (R-524), not a fault.
@@ -0,0 +1,76 @@
# 07 — SCENARIO C: the CONTROLLER-ONLY restart during an app update (wishlist v0.66.0 -> v0.67.0)
#
# Run by the main session on 2026-09-21. This is the scenario that matters most for R-608: it is
# EXACTLY what a controller self-update does to a running app update — the controller container
# restarts, the app's own containers keep running. Guest 9202, controller v0.260.0 (the lock is NOT
# in this build; that is deliberate — this measures the behaviour the lock is meant to make moot).
== TIMESTAMP TABLE ==
poll cadence : 200 ms
app confirmed healthy : 12:35:53.969 UTC (state=running health=True)
Update pressed : 12:35:53.971 UTC -> 202 "Frissites elindult"
phase AT THE DECISION : "starting" observed 12:36:33.136 UTC
cut command : ssh demo-hp 'pct exec 9202 -- systemctl restart felhom-controller-bootstrap.service'
cut command LATENCY : 1 675 ms (rc=0)
phase the box RECORDED : "verifying" (from the controller's own recovery line)
-> THE SAME INSTRUMENT FINDING AS A AND B, REPRODUCED WITH A DIFFERENT AND MUCH FASTER CUT.
The decision was taken at `starting` and the box still recorded `verifying`. A controller
restart returns in 1.7 s where `pct stop` took 3.0-3.8 s, and `starting` STILL could not be
caught. Measured on this box, `starting` lasts well under a second for these apps.
**`RecoverUpdates` handles `starting` and `verifying` in ONE branch (`update.go:909`), so all
three scenarios exercise the same recovery arm** — the arm that says "the new version may have
run". That is the arm the brief wanted measured, and it is measured three times.
== THE APP'S OWN CONTAINERS DURING THE CONTROLLER RESTART ==
felhom-controller 0.260.0 Up 1 second (health: starting) <- restarted
wishlist ghcr.io/cmintey/wishlist:v0.67.0 Up 1 second (health: starting) <- the update's own `up -d`
uptime-kuma louislam/uptime-kuma:2.5.0 Up 3 minutes (healthy) <- UNDISTURBED
vikunja vikunja/vikunja:2.6.0 Up 3 minutes <- UNDISTURBED
glance glanceapp/glance:v0.8.5 Up 3 minutes (healthy) <- UNDISTURBED
filebrowser gtstef/filebrowser:1.3.3-stable Up 3 minutes (healthy) <- UNDISTURBED
-> The controller restart does NOT restart the apps. Only the app the update itself was recreating
shows a new uptime, and that is the update's `up -d`, not the restart.
== THE RECOVERY, verbatim ==
2026/09/21 12:36:34 update.go:909: [WARN] [stacks] update recovery: wishlist was interrupted in verifying (started 2026-09-21T12:35:54Z) — the new version may have run; marking it Updating and RESUMING the health wait
2026/09/21 12:36:34 update.go:951: [INFO] [stacks] update wishlist: resuming after a controller restart — `up -d` then the health wait
2026/09/21 12:36:35 update.go:858: [INFO] [stacks] update wishlist: phase verifying
2026/09/21 12:36:45 update.go:652: [INFO] [stacks] update wishlist: healthy after 10s (the app's health check passed)
2026/09/21 12:36:45 update.go:658: [INFO] [stacks] update wishlist: DONE in 51s
== THE FOUR VERSION OBSERVABLES, SIDE BY SIDE ==
pinned_images : {"wishlist": "ghcr.io/cmintey/wishlist:v0.67.0"}
installed_images : {"wishlist": "ghcr.io/cmintey/wishlist:v0.67.0"}
live compose line: image: ghcr.io/cmintey/wishlist:v0.67.0
docker inspect : ghcr.io/cmintey/wishlist:v0.67.0 | running
-> ALL FOUR AGREE.
== END STATE, AND WHAT THE HOUSEHOLD SEES ==
state=running updating=False update_phase=done label='Frissitve'
update_error=(none) hold_reason=(none) health=True
update-journal.json : ABSENT (cleared)
app page, Hungarian : <span class="tag tag-ok" title="Ez az alkalmazas a legfrissebb elerheto valtozatot futtatja.">Naprakesz</span>
app page, English : <span class="tag tag-ok" title="This app is running the newest version available.">Up to date</span>
(accents transliterated here only to keep this file ASCII-searchable; intact in the page)
no `data-update-error` block on either page — there is nothing to apologise for.
SEARCH CONTROLS: POSITIVE 'Naprak' in hu = 1 · POSITIVE 'Up to date' in en = 1 ·
NEGATIVE 'ZZZ-not-present' = 0
== STOP CONDITIONS — none tripped ==
updating stuck true : NO journal surviving : NO pin vs running image : AGREE
retry loop : NO hold with no sentence : N/A (no hold)
data : wishlist's seeded list item was NOT re-read by this session — see the gap note below.
== GAP, STATED ==
The seeded-data read-back through wishlist's front door was not repeated after this scenario. The
measuring agent seeded and read it back BEFORE the drill (02-seeding.txt) and the app is healthy
and serving on v0.67.0 after it, but "the data is still there" is NOT asserted here for C. A and B
both carry a real post-cut read-back; C does not.
== VERDICT ==
A controller-only restart during an app update ends HONEST. The update resumes, completes, the
other apps are untouched, and the household is shown a clean „Naprakesz" with no error. This is
the behaviour v0.261.0's lock makes unnecessary rather than fixes — worth knowing, because it
means the lock is defence in depth, not a repair of something broken.
@@ -0,0 +1,48 @@
# 08 — THE v0.261.0 LOCK, PROVEN LIVE (R-608), with a positive AND a negative control
#
# Guest 9202, controller v0.261.0, 2026-09-21. The question is not "is the callback wired" but
# "does the swap actually refuse, and is it OUR refusal doing it".
== WHY A CONTROL WAS NEEDED, AND WHY THIS ONE IS THE RIGHT ONE ==
Guest 9202 has NO host agent wired, so `TriggerUpdate` refuses anyway — with the AGENT's sentence.
That makes it the perfect negative control: in `TriggerUpdate` the new app-update check sits BEFORE
the agent check, so if the lock fires we see OUR sentence and if it does not we see the AGENT's.
Two different sentences, one probe. It also makes the probe SAFE: no swap can reach a machine.
(Self-update is `enabled: false` on this guest by design. It was turned on for this probe and
turned back OFF immediately afterwards — the original file was copied to
/root/controller.yaml.pre-lock-probe first and restored from it. Verified reverted below.)
== CONTROL A — self-update status, no app update running ==
GET /api/selfupdate/status -> {"ok":true,"data":{"running":false}}
== CONTROL B (NEGATIVE) — the MANUAL trigger with NO app update running ==
12:39:55.917 POST /api/selfupdate/update
-> {"ok":false,"error":"A frissites nem erheto el (nincs gazda-ugynok)"}
i.e. the AGENT refusal. The lock is NOT firing, because nothing is in flight. Correct.
== THE PROBE (POSITIVE) — the MANUAL trigger WHILE an app update is in flight ==
12:39:55.941 POST /api/stacks/uptime-kuma/update -> 202 "Frissites elindult"
12:39:56.241 app update IN FLIGHT, phase = safety-dump
12:39:56.244 POST /api/selfupdate/update
-> {"ok":false,"error":"Egy alkalmazas frissitese eppen folyamatban van. A vezerlo frissitese utana inditható."}
THE SENTENCE CHANGED. That is the v0.261.0 gate firing, and firing BEFORE the agent check.
12:39:56.270 GET /api/selfupdate/status -> {"running":false} — nothing was started.
(accents transliterated here only to keep this file ASCII-searchable; intact on the wire)
== WHAT THE PHASE TELLS US, and it is worth stating ==
The probe landed in `safety-dump`. That phase is NOT covered by the pre-existing `backupRunning`
gate — only `backing-up` is, because that one takes the backup single-flight. So the probe hit
exactly the window the new lock exists for, rather than a window that was already protected.
== THE REVERSE DIRECTION — NOT staged live, and why ==
`UpdatePreflight` refusing an app update with reason `self_updating` while the controller swaps is
covered by `TestR608_PreflightRefusesWhileTheControllerSwaps` (a consequence test, red-proofed by
deleting the block). Staging it LIVE would need a real controller swap in flight, which needs a host
agent this guest does not have and a genuine image swap mid-measurement. **Not measured live. Stated
as a gap rather than implied.**
== CONFIG RESTORED ==
self_update.enabled is back to `false` on guest 9202, verified by re-reading the file after the
restore, and the controller was restarted on the restored config.
@@ -0,0 +1,12 @@
2026-09-21T12:55:11Z === pass 1/3
2026-09-21T12:55:11Z glance BEHIND and within a major — pressing Update
2026-09-21T12:55:11Z glance REFUSED reason=downgrade TERMINAL — will not press again. Ez a változat újabb a katalógusban lévőnél — visszalépés csak az üzemeltető kérésére.
2026-09-21T12:55:11Z uptime-kuma BEHIND and within a major — pressing Update
2026-09-21T12:55:11Z uptime-kuma REFUSED reason=downgrade TERMINAL — will not press again. Ez a változat újabb a katalógusban lévőnél — visszalépés csak az üzemeltető kérésére.
2026-09-21T12:55:11Z vikunja BEHIND and within a major — pressing Update
2026-09-21T12:55:11Z vikunja REFUSED reason=downgrade TERMINAL — will not press again. Ez a változat újabb a katalógusban lévőnél — visszalépés csak az üzemeltető kérésére.
2026-09-21T12:55:11Z wishlist BEHIND and within a major — pressing Update
2026-09-21T12:55:11Z wishlist REFUSED reason=downgrade TERMINAL — will not press again. Ez a változat újabb a katalógusban lévőnél — visszalépés csak az üzemeltető kérésére.
2026-09-21T12:56:11Z === pass 2/3
2026-09-21T12:57:11Z === pass 3/3
2026-09-21T12:57:12Z === summary outcomes={} never_again=['glance', 'uptime-kuma', 'vikunja', 'wishlist']
@@ -0,0 +1,75 @@
# 09 — THE UNATTENDED NIGHT: an app updates with nobody pressing anything
#
# This is the phase the previous session's brief contained and **did not run, and did not say so**
# (R-611). It runs here. The caller is `unattended-caller.py` in this directory — EVIDENCE, NOT
# PRODUCT: it presses exactly `POST /api/stacks/<n>/update`, the same button a person presses, and
# nothing else. No controller code was added for it. Guest 9202, controller v0.261.0.
## Scenario F — the success night: PROVEN
`uptime-kuma` **2.5.0 → 2.5.1** was applied by the caller with no human action, against a real
one-step catalog edge (DRILL 2, `ae08a037fd68`). Confirmed on the box afterwards:
uptime-kuma installed = louislam/uptime-kuma:2.5.1 catalog = louislam/uptime-kuma:2.5.1
updating = False update_phase = done hold = none
**HONEST GAP, and it is the coordinator's own instrumentation error:** that first run's stdout was
piped through `tail`, which buffers, and the run was later killed — **so the caller's own log lines
for the F press were lost.** The OUTCOME is solid (the box moved 2.5.0 → 2.5.1 and only the caller
pressed it), but the per-pass log for F is gone. The second run below was written straight to a file
for exactly this reason. Recorded rather than quietly omitted.
## Scenario G — two results, and the first one is the more interesting
### G-a. The deliberately broken edge was NEVER ATTEMPTED — the safety rule filtered it first
The C3-class negative control was `vikunja: 2.6.0 → alpine:3.20`. The caller **skipped it**, because
`alpine:3.20` and `vikunja/vikunja:2.6.0` are different repositories and therefore cannot be ordered,
so the edge is "across" and belongs to a human (`09` §3 decision 3, §3b Q3). The box was never asked
to run it: `vikunja` ended the night still on 2.6.0, no update, no hold, untouched.
**That is a real finding for Slice 6 and it cuts both ways.** The within-a-major rule is the FIRST
line of defence and it works — a catalog edge that cannot be ordered never reaches the guarded update
unattended. But it also means **this shape of broken edge cannot be used to measure the unattended
HOLD path**, because the rule that makes automatic updates safe is the same rule that refuses it.
### G-b. The no-retry property, PROVEN over three passes
Reverting the catalog (`f5f6a152b513`) left all four apps running something NEWER than the catalog —
R-524's Ahead state. The caller then had a genuine TERMINAL refusal to react to:
pass 1/3 glance BEHIND and within a major — pressing Update
glance REFUSED reason=downgrade TERMINAL — will not press again.
„Ez a változat újabb a katalógusban lévőnél — visszalépés csak az
üzemeltető kérésére."
... the same for uptime-kuma, vikunja and wishlist ...
pass 2/3 (nothing)
pass 3/3 (nothing)
summary outcomes={} never_again=['glance','uptime-kuma','vikunja','wishlist']
**Four apps, pressed exactly once each, then never again across two further passes.** Nothing on the
box changed: no update ran, no hold was set, no journal was written.
**This is R-524 and R-609 working together end to end, unattended.** Before R-609 put `reason` on the
wire the caller would have had only a Hungarian sentence to parse, and the only safe readings were
"give up on everything" or "press for ever". The full log is `09-unattended-night.log`.
## What Slice 6's design now knows that it did not
| question | answer, from measurement |
|---|---|
| how long does one app take end to end, unattended? | 51 s – 1 m 26 s for these four small apps, including the health wait — see 04/05/07 |
| does the caller need new controller code? | **No.** It presses the existing guarded Update and reads the existing state. |
| can it tell "wait" from "never"? | **Yes, since v0.261.0** — and not before |
| does the within-a-major rule hold? | **Yes, and it is the first thing that fires** — G-a |
| what does a held app look like the next morning? | **STILL UNKNOWN unattended** — see below |
## NOT MEASURED, stated plainly
**The unattended HOLD path.** `09` §3b Q4 asks who is told when an automatic update ends HELD. This
night never produced a hold, because the only failing edge available was one the safety rule
correctly refused (G-a). Measuring it needs an edge that **passes** the within-a-major test and still
fails its health check — same repository, same major version, a tag that starts and does not serve.
Real images rarely offer one, so this probably needs a purpose-built image rather than a catalog
move. **Q4 therefore still rests on the ATTENDED hold measured in slice 4 (v0.238.0, Scenario F),
not on an unattended one.** Recorded as the gap it is.
@@ -0,0 +1,167 @@
#!/usr/bin/env python3
"""unattended-caller.py — the Slice 6 spike: an app updates with NOBODY pressing anything.
THIS IS EVIDENCE, NOT PRODUCT. It lives under documentation/audits/ and nothing in the controller
imports it. It exists to MEASURE the mechanism Slice 6 would need before that slice is designed, by
pressing exactly the same guarded Update a person presses — `POST /api/stacks/<n>/update` — and
nothing else. No new endpoint, no new controller code, no privileged path.
WHAT IT DOES NOT DO, deliberately: it does not decide policy. The window is simulated by running it;
the per-app switch of `09` §3b Q2 does not exist yet; it never touches a box that is not the scratch
guest. Run it from DooPlex against guest 9202 only.
THE TWO RULES IT EXISTS TO PROVE
1. within a major, for EVERY compose service, or it does not press at all (§3 decision 3, §3b Q3);
2. a refusal's REASON decides whether it ever presses again (R-609):
TRANSIENT busy updating deploying migrating self_updating -> try next pass
TERMINAL held downgrade -> never again
FOR A HUMAN memory disk no_backup -> log and leave alone
Before R-609 the body carried only a Hungarian sentence, so this distinction was unavailable to
anything that is not a person — which is why a caller like this could not have been written.
Usage: python3 unattended-caller.py --passes 6 --every 300
"""
import argparse, json, re, subprocess, sys, time
CALL = "/tmp/ctl/c.sh" # the helper from 00-api-recipe.md
TRANSIENT = {"busy", "updating", "deploying", "migrating", "self_updating"}
TERMINAL = {"held", "downgrade"}
FOR_A_HUMAN = {"memory", "disk", "no_backup"}
TAG = re.compile(r"^v?(\d+)(?:\.(\d+))?(?:\.(\d+))?(.*)$")
def log(msg):
print("%s %s" % (time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()), msg), flush=True)
def call(method, path, data=None):
cmd = [CALL, method, path] + ([data] if data else [])
out = subprocess.run(cmd, capture_output=True, text=True, timeout=180).stdout
try:
return json.loads(out)
except Exception:
return {"_raw": out}
def split_ref(ref):
"""(repo, tag) or (None, None) when the reference carries no plain tag."""
if "@" in ref:
return None, None
i = ref.rfind(":")
if i < 0 or "/" in ref[i + 1:]:
return None, None
return ref[:i], ref[i + 1:]
def same_major(a, b):
"""True only when BOTH tags are plain versions, share a suffix, and share a first number.
Mirrors stacks.CompareImageRefs deliberately: the box's own rule is the one under test, and a
caller that judged 'within a major' differently would measure its own opinion instead.
"""
ra, ta = split_ref(a)
rb, tb = split_ref(b)
if ra is None or ra != rb:
return False
ma, mb = TAG.match(ta or ""), TAG.match(tb or "")
if not ma or not mb:
return False
if ma.group(4) != mb.group(4): # the suffix must be IDENTICAL (…-apache vs …-apache)
return False
if ma.group(2) is None or mb.group(2) is None:
return False # one component is a LINE, not a version
return ma.group(1) == mb.group(1)
def edge_is_within_a_major(st):
"""ALL services must pass. One unorderable service makes the whole edge 'across' -> a human."""
installed = {k: v["ref"] for k, v in (st.get("app_config", {}).get("installed_images") or {}).items()}
catalog = st.get("catalog_images") or {}
if not installed or not catalog or set(installed) != set(catalog):
return False, "service sets differ or nothing recorded"
for svc, want in catalog.items():
got = installed[svc]
if got == want:
continue
if not same_major(got, want):
return False, "across: %s %s -> %s" % (svc, got, want)
return True, "within a major"
def follow(name, timeout=900):
"""Watch one update to its end. Returns done | failed | held | timeout."""
deadline = time.time() + timeout
last = None
while time.time() < deadline:
st = call("GET", "/api/stacks/%s" % name)
phase, updating = st.get("update_phase"), st.get("updating")
if phase != last:
log(" %s: phase=%s updating=%s" % (name, phase, updating))
last = phase
if not updating and phase in ("done", "failed"):
held = bool(st.get("hold_reason"))
return "held" if held else phase
time.sleep(2)
return "timeout"
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--passes", type=int, default=6)
ap.add_argument("--every", type=int, default=300)
args = ap.parse_args()
never_again, outcomes = set(), {}
for p in range(1, args.passes + 1):
log("=== pass %d/%d" % (p, args.passes))
call("POST", "/api/sync")
call("POST", "/api/stacks/rescan") # R-607: a sync can say 'no change' and still move
stacks = call("GET", "/api/stacks")
stacks = stacks if isinstance(stacks, list) else stacks.get("data", [])
for st in stacks:
name = st.get("name")
if not st.get("deployed") or st.get("protected") or name in never_again:
continue
installed = {k: v["ref"] for k, v in (st.get("app_config", {}).get("installed_images") or {}).items()}
if not installed or installed == (st.get("catalog_images") or {}):
continue # unknown, or level with the catalog
ok, why = edge_is_within_a_major(st)
if not ok:
log(" %s SKIP — %s" % (name, why))
continue
log(" %s BEHIND and within a major — pressing Update" % name)
r = call("POST", "/api/stacks/%s/update" % name)
if r.get("ok") is False:
reason = (r.get("data") or {}).get("reason", "")
sentence = r.get("error", "")
if reason in TERMINAL:
never_again.add(name)
log(" %s REFUSED reason=%s TERMINAL — will not press again. %s" % (name, reason, sentence))
elif reason in TRANSIENT:
log(" %s REFUSED reason=%s transient — retry next pass. %s" % (name, reason, sentence))
elif reason in FOR_A_HUMAN:
never_again.add(name)
log(" %s REFUSED reason=%s — needs a person. %s" % (name, reason, sentence))
else:
never_again.add(name)
log(" %s REFUSED reason=%r UNKNOWN — stopping on it, safest reading. %s" % (name, reason, sentence))
continue
started = time.time()
end = follow(name)
outcomes[name] = (end, round(time.time() - started, 1))
log(" %s ENDED %s after %.1fs" % (name, end, time.time() - started))
if end in ("held", "timeout"):
never_again.add(name)
if p < args.passes:
time.sleep(args.every)
log("=== summary outcomes=%s never_again=%s" % (outcomes, sorted(never_again)))
return 0
if __name__ == "__main__":
sys.exit(main())
+8 -1
View File
@@ -714,7 +714,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-516** | **[P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item **„Debug"**; (2) the dashboard CPU tile **„Load: 0.29 / 0.39 / 0.37"**; (3) the launcher tile **„Filebrowser"** opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) **Uptime Kuma 2.4 opens on „Which database would you like to use?"** (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), `wiki.DOMAIN` (R-498). **Fix shape:** rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed `db-config.json` for SQLite in the template (catalog). **Added by F4 (20:04:54Z):** (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. **Added by F7 (20:50–21:00Z, system disk at 95 %):** (10) a banner on every page in English, „**SSD disk usage high: 90%**"; (11) the dashboard tile reads „Rendszer (/) 61.8 GB / 68.7 GB (**90%**)" while `df` reports **95 %** (reserved blocks ignored), and „(/)" labels the data volume `/mnt/sys_drive`; the deploy page says nothing about free disk. **Added by the i18n spike (2026-09-17, controller v0.247.0):** (12) **six formal („ön") forms in the converted dashboard copy** — „Olvassa be telefonnal" (launcher QR hint), „Biztosan kikapcsolja a megosztást?" (launcher), „Ha újratelepíti … importálnia kell" (the layout's remove-app modal), „Kérjük, vegye fel a kapcsolatot" (backups empty state). They were NOT fixed — a localisation release may not change Hungarian bytes — and are now COUNTED by `controller/scripts/i18n_missing_gate.py` (`HU_FORMAL_CEILING` = 6, a ratchet: a new one convicts, fixing one lowers it). The inventory (`audits/I18N-INVENTORY-2026-09-17.md`) is the list this row closes against in localisation slice 6 (R-561). **Extended 2026-09-17 by slice 1 (controller v0.248.0–v0.250.0):** the converted copy now counts **16** formal forms (`HU_FORMAL_CEILING` 6 → 10 → 12 → 16, each raise stated, Hungarian unchanged by rule) — and the count is an UNDER-count: the gate's stem list sees 4 on the release C pages (storage, drive wizards, debug), while a wider list („írja be”, „adja meg”, „adjon hozzá”, „válassza ki”, „biztosan eltávolítja”, „engedélyezze”, „hozzon létre”, „kattintson”) finds at least **22** keys there alone (login „Adja meg a jelszavát”, the NAS guide, the storage confirms). Widen the stems when this row is worked; the ceiling rises with them. **NARROWED 2026-09-20 by localisation slice 6's walk, item by item** (`audits/i18n-slice6-2026-09-20/R-516-item-by-item.md`): item 2 FIXED (slice 1), items 3 and 6 APP-OWNED (and for an English household item 6 is an advantage), item 11 is a wrong NUMBER rather than a language and needs its own row, item 1 is correct English for an English reader and still open for a Hungarian one. **Items 4, 7, 8, 9 and 10 are NOT DECIDABLE FROM AN ENGLISH WALK** - they are about Hungarian copy quality, a disconnected drive and a full disk, none of which this walk had. **What this row is now waiting for is a HUNGARIAN walk on a box with a second drive**, not another English one; saying it closed would be the kind of closure that makes a register stop meaning anything. | **NARROWED 2026-09-20 - rank P3-LOW; owner: CC; needs a Hungarian walk** |
| **R-518** | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — rank P2-MEDIUM; owner: CC (controller)** |
| **R-519** | **[P2-MEDIUM] After a backup torn by a power cut, an app's restore point carries the new database dump's time while its files are from the previous run — and no customer screen says the run was interrupted.** MEASURED 2026-09-14 (BIGNIGHT F2, VM 333): „Mentés most" 19:40:03Z; the power was cut 19:40:08Z while adventurelog was stopped for its volume dump. On disk afterwards, `backups/primary/adventurelog`: `db-dumps/adventurelog-postgres.sql` **19:40:07**, `volume-dumps/*.tar` **19:00:35**, `manifest.json created_at 19:02:34Z`; bookstack the same shape (sql 19:40:07, tars 19:00:48). `GET /api/backup/snapshots` for both → `time 2026-09-14T19:40:07Z, helyi` (`phase5/F2/units-on-disk.txt`, `backup-honesty.txt`). `/backups/apps` shows „Utolsó adatbázis mentés 2026-09-14 21:40 · … OK" and every app „Utolsó: 5 perce"; `/backups` and `/dashboard` contain no word of an interruption (fragments `megszakad|sikertelen|nem sikerült` = 0; control „hiba" appears in the standing warning text). The controller itself knew: `[appstop] crash recovery: an app-data backup (volume dump) … was interrupted … restarting them: [adventurelog]` and pushed `backup_failed (error)` to the hub. A household restoring „the 21:40 backup" gets 21:00 files for BookStack's uploads. **Fix shape:** date a point by the oldest part it contains (or mark it partial) and show the interrupted run on the backups page until the next complete one. **F6 (drive unplugged 1 s into a backup, 20:37:59Z) adds three facts:** the run skipped four apps' volume dumps („Skipping volume dump for immich — drive disconnected", also jellyfin, nextcloud, paperless-ngx) and still reported `db_dump {"count":6, "success":true}`; nextcloud's point is dated 20:37:59Z (its SQL finished before the unplug) beside 19:01 volume tars; and a torn `immich-postgres.sql.tmp` (20:38:00) plus F3's `pre-restore-…-nextcloud-mariadb.sql.tmp` are left in the units on the drive. Immich's point correctly stayed at 19:02:34Z (the `.tmp` was not promoted). `/backups/apps` fragments `kihagy|sikertelen|részleges` = 0 (`phase5/F6/`). | **READY — rank P2-MEDIUM; owner: CC (controller)** |
| **R-520** | **[P3-LOW] A power cut during a guarded Update leaves no record that an update was running, so whether the journal resumes or aborts honestly is unmeasured for a real version change.** MEASURED 2026-09-14 (BIGNIGHT F3, VM 333): the only Update the catalog allowed after the Phase 4 revert was a same-version one on nextcloud; power was cut 2 s after `update nextcloud: phase pulling` (safety dump written, pin advanced to the unchanged definition). After boot: no log line resumes, aborts or names the interrupted update; `app.yaml` `pinned_images` = `installed_images`, no hold, no verdict; page „Fut · Naprakész". Nothing wrong was produced — and nothing could have been, with identical images. **What it needs:** the same cut on a real bump (a throwaway app with a one-step catalog move on a scratch branch or the scratch guest), asserting the page's message and the pin after boot. **CLOSED 2026-09-21 — measured on a REAL bump, and it behaved.** Scratch guest 9202 (demo-hp), controller v0.260.0, throwaway `uptime-kuma`, real catalog move 2.4.0 -> 2.5.0 (reverted the same session). `pct stop 9202` at 11:04:54.439Z, 3 ms after the poller observed `update_phase: pulling`. **Honest limit on the instrument:** `pct stop` returned 3.8 s later, so the phase at the DECISION is observed and the phase at the freeze is inferred. **After boot the box said so itself — a POSITIVE observable, not an absence:** `update recovery: uptime-kuma was interrupted in pulling (started 2026-09-21T11:04:52Z) — nothing had run; putting the pin back`, then `pin uptime-kuma: …:2.4.0` and `update …: pin and definition PUT BACK to the pre-update version`. `pinned_images` and `installed_images` both 2.4.0, live compose line 2.4.0, no hold (correctly — nothing ran), `update_phase: failed`, app **running and healthy on 2.4.0** 31 s after boot, and the household told: „A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább." **THE INSTRUMENT TRAP THIS RUN FOUND IS WORTH MORE THAN THE RESULT.** The first post-crash read, taken from the host against the STOPPED guest, reported the update journal **ABSENT** — which would have made this a defect finding. It was a FALSE NEGATIVE: **`pct mount` maps the guest's ROOTFS ONLY and does not apply the guest's own internal mounts**, and 9202 keeps `/var/lib/docker` on its own ext4 mount under the `mp0` volume, so from the host that path is an EMPTY STUB. The bytes were at `/var/lib/lxc/9202/rootfs/var/lib/felhom/docker/volumes/felhom-controller-data/_data/data/`, and the box's own recovery read them at the next boot. Caught by running `find` over the whole rootfs instead of trusting one constructed path, and by demanding a positive control for the directory searched. **An empty directory is not evidence of an absent file.** **What is NOT measured and does not re-open this row:** the second cut, in the `starting` phase (Scenario F's territory — the migration may have run). Evidence: `audits/update-arc-2026-09-21/` 08-12. | **CLOSED 2026-09-21 — measured on a real version change** |
| **R-520** | **[P3-LOW] A power cut during a guarded Update leaves no record that an update was running, so whether the journal resumes or aborts honestly is unmeasured for a real version change.** MEASURED 2026-09-14 (BIGNIGHT F3, VM 333): the only Update the catalog allowed after the Phase 4 revert was a same-version one on nextcloud; power was cut 2 s after `update nextcloud: phase pulling` (safety dump written, pin advanced to the unchanged definition). After boot: no log line resumes, aborts or names the interrupted update; `app.yaml` `pinned_images` = `installed_images`, no hold, no verdict; page „Fut · Naprakész". Nothing wrong was produced — and nothing could have been, with identical images. **What it needs:** the same cut on a real bump (a throwaway app with a one-step catalog move on a scratch branch or the scratch guest), asserting the page's message and the pin after boot. **CLOSED 2026-09-21 — measured on a REAL bump, and it behaved.** Scratch guest 9202 (demo-hp), controller v0.260.0, throwaway `uptime-kuma`, real catalog move 2.4.0 -> 2.5.0 (reverted the same session). `pct stop 9202` at 11:04:54.439Z, 3 ms after the poller observed `update_phase: pulling`. **Honest limit on the instrument:** `pct stop` returned 3.8 s later, so the phase at the DECISION is observed and the phase at the freeze is inferred. **After boot the box said so itself — a POSITIVE observable, not an absence:** `update recovery: uptime-kuma was interrupted in pulling (started 2026-09-21T11:04:52Z) — nothing had run; putting the pin back`, then `pin uptime-kuma: …:2.4.0` and `update …: pin and definition PUT BACK to the pre-update version`. `pinned_images` and `installed_images` both 2.4.0, live compose line 2.4.0, no hold (correctly — nothing ran), `update_phase: failed`, app **running and healthy on 2.4.0** 31 s after boot, and the household told: „A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább." **THE INSTRUMENT TRAP THIS RUN FOUND IS WORTH MORE THAN THE RESULT.** The first post-crash read, taken from the host against the STOPPED guest, reported the update journal **ABSENT** — which would have made this a defect finding. It was a FALSE NEGATIVE: **`pct mount` maps the guest's ROOTFS ONLY and does not apply the guest's own internal mounts**, and 9202 keeps `/var/lib/docker` on its own ext4 mount under the `mp0` volume, so from the host that path is an EMPTY STUB. The bytes were at `/var/lib/lxc/9202/rootfs/var/lib/felhom/docker/volumes/felhom-controller-data/_data/data/`, and the box's own recovery read them at the next boot. Caught by running `find` over the whole rootfs instead of trusting one constructed path, and by demanding a positive control for the directory searched. **An empty directory is not evidence of an absent file.** **What is NOT measured and does not re-open this row:** the second cut, in the `starting` phase (Scenario F's territory — the migration may have run). Evidence: `audits/update-arc-2026-09-21/` 08-12. **CORRECTED 2026-09-21 (same day): that sentence left the dangerous half of the question in prose and in no row — see R-610, which carries it and CLOSES it with three measurements.** The history above stands as written; only this pointer is added. | **CLOSED 2026-09-21 — measured on a real version change** |
| **R-521** | **[P3-LOW] One unplugged drive sends the operator five e-mails and the household none.** MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): `storage_disconnected (error)` at 21:58:02 CEST plus `app_start_failed (warning)` for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five `Operator email sent`). The customer's mailbox (`tester1@felhom.eu`, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. **Fix shape:** suppress `app_start_failed` for apps stopped by a `storage_disconnected` (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. **F6, 40 min later, the opposite failure:** a second, separate drive loss (20:38:33Z) produced `storage_disconnected (error)` and four `app_start_failed`, and the hub logged `Operator email suppressed … cooldown` for all five — **no mail at all for the second unplug**; only `health_degraded (warning)` mailed. A per-key cooldown that outlives the recovery (`storage_reconnected` came between them) silences a new incident. **F7:** the system disk at 95 % produced only `health_degraded (warning)`, whose operator mail was **suppressed by the cooldown** left by F6's `health_degraded` 15 minutes earlier; no disk-specific event reached the hub at all — the operator was not told the disk was nearly full. | **READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy)** |
| **R-522** | **[P3-LOW] While the box has no internet, the dashboard's „Cloudflare Tunnel" tile keeps saying „Fut", and no page tells the household the box is offline.** MEASURED 2026-09-14 (BIGNIGHT F8, VM 333): VM 333's traffic off the LAN and to the hub was dropped at demo-hp's bridge 21:08:36 → 21:26:07Z. Throughout, the LAN dashboard (probed every 26 s from demo-hp) answered 200 and, polled every 2 min, showed no banner and the tile „Cloudflare Tunnel — Biztonságos internetkapcsolat — a szerver portnyitás nélkül érhető el kívülről. · **Fut** · Védett"; meanwhile cloudflared logged ≈ 20 errors every 2 minutes, the public name answered 530, and the controller logged `[report] Push failed … context deadline exceeded` and `Job hub-report failed: hub push failed after 3 attempts`. The tile reports the container, not the connection. A household whose remote access is gone sees „Fut". **Fix shape:** the tile reads the tunnel's connection state (cloudflared's registered connections or the report push result) and says „Nincs internetkapcsolat" when either fails. | **READY — rank P3-LOW; owner: CC (controller)** |
| **R-524** | **[P2-MEDIUM] When the catalog moves an app back to an older version, a box that already updated shows „Frissítés elérhető" — and the offered Update is a downgrade.** MEASURED 2026-09-15 (BIGNIGHT Phase 6, VM 333): privatebin was updated 2.0.5 → 2.0.6 through the guarded Update after the drill bump; the catalog was then reverted to 2.0.5 (`a161ccb`). At 22:13:37Z the box reads `installed privatebin/pdo:2.0.6`, `catalog privatebin/pdo:2.0.5`, `catalog_since 2026-09-14`, and the app page tag „**Frissítés elérhető — ma**" with the title „Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot." The label compares for *difference*, not for *newer* (`09-update-architecture.md` §5.4 render table); the guarded Update would advance the pin „to the catalog's current definition" — 2.0.6 → 2.0.5. The same state follows any real upstream yank. **Not pressed tonight.** **Fix shape:** compare versions (or `catalog_since` against the installed record) and render „Naprakész" / „a katalógusnál újabb" when the box is ahead; refuse a pin move to an older tag without an operator word. **CLOSED 2026-09-21 — controller v0.260.0.** `stacks.CatalogOrder` (`internal/stacks/updateorder.go`) replaces the three-way comparison with FOUR verdicts — Unknown / Current / Behind / **Ahead** — and **moves out of `web` so the badge and the refusal read ONE verdict**; `web.compareInstalledToTemplate` is now a thin wrapper. An app AHEAD reads „Naprakész" / "Up to date" with `tag-ok` (the same word and class as level — there is nothing for the household to do) and a title saying why (`badge.update.ahead.title`, both bundles). `Manager.UpdatePreflight` refuses with reason `downgrade`, HTTP 409, „Ez a változat újabb a katalógusban lévőnél — visszalépés csak az üzemeltető kérésére.", logged with both image maps. **The API now renders update refusals through `errText`**, so the new key is not a seam built and never wired. **Ahead is NARROW on purpose:** every differing service must be orderable AND newer, or the verdict falls back to Behind — this gate can BLOCK an update, so it errs towards letting one run. Ordering is `util.Version.Compare` (the house rule: one comparator) behind a tag normaliser — `X.Y`/`X.Y.Z`, optional leading `v`, two-part padded with `.0`, and a trailing suffix that must be IDENTICAL on both sides, so `nextcloud:31.0.14-apache → 31.0.15-apache` orders while `postgres:16-alpine`, `26.05.2-ls310 → -ls311`, `kimai/kimai2:apache-2.57.0`, a date stamp and a digest pin do not. **The suffix rule was found by the fixture, not by design** — the first implementation called every real catalog tag unorderable. **Three red-proofs, each SEEN to fail.** Recorded as `09` §3 decision 10 (decided by CC unattended — operator may reverse). | **CLOSED 2026-09-21 — controller v0.260.0** |
@@ -781,6 +781,13 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-605** | **[P3-LOW] A catalog gate that REFUSED TO RUN and a gate that ran and could not decide print the same word, so a reader cannot tell which happened.** FOUND 2026-09-21 while answering why the chaos night's update round could not run. On 2026-09-17 `check-image-resolvable` and `check-volume-persistence` both returned INCONCLUSIVE and the drawn `update` action was replaced with `use` (`audits/DRILL-chaos-night-2026-09-17.md:692-695`). **Neither script is defective — they behaved exactly as designed**, and both headers say why: a detector that cannot prove itself must refuse to report rather than guess (`check-image-resolvable.py` cites the 2026-07-21 incident where a Docker Hub throttle read as 24 of 65 pins falsely dead). **What is missing is the DISTINCTION.** `check-volume-persistence.py`'s `self_test` refuses to evaluate ANY app when it cannot build its canary image — a harness-level refusal — while `classify()` returns a per-app UNDETERMINED for an app that wrote nothing; `check-image-resolvable.py` likewise separates a harness-level canary failure (exit 2 at `check()` L180-183) from a per-pin throttle (L121-128). **`catalog_gates.py`'s VERDICT map collapses all of them into one `INCONCLUSIVE` label**, so the operator-facing summary cannot say whether the gate ran at all. **The cost is real and already paid:** no raw stdout of the 2026-09-17 run survives in either evidence directory, so the exact triggering path is INFERRED from the code plus the documented throttle precedent, not observed — a distinct summary line would have recorded it for free. **Fix shape:** have each gate's exit distinguish "the harness refused" from "the result is undetermined" (a third exit code, or a marker line the runner matches), and have `catalog_gates.py` print the two differently. **Ships with a decoy each way (R-421): a run whose canary fails must NOT read as a per-app undetermined, and vice versa.** Small. | **READY — rank P3-LOW; owner: CC (catalog)** |
| **R-606** | **[P2-MEDIUM] Every sentence the UPDATE path shows a household is Hungarian-only, and it lands on TWO pages — FOUND LIVE 2026-09-21, not by reading.** While measuring R-520 on guest 9202 at controller v0.260.0, the English app page rendered the badge in English — *"Update available — today"* — directly above *„A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább."* **The mixed line is worse than either language whole**, and this one is the household's only explanation for why their app did not move. **MECHANISM:** `Stack.UpdateError` is a finished Hungarian STRING, not a key. `Manager.finishUpdate` stores it (`internal/stacks/update.go` L517, 526, 530, 679, 902, 907, 936, 943) from the raw literals `MsgUpdateInterrupted`, `MsgUpdatePullFailed`, `MsgUpdateBackupFailFmt`, `MsgUpdateDumpFailFmt`, `MsgUpdatePinFailed`, `MsgUpdateJournalFailed`, `MsgUpdateBackupNoUnit`, `MsgUpdateHoldUnsaved` (`update.go` L79-96), and BOTH `app_info.html:36` and `stacks.html:99,104` render it verbatim. **SCOPE IS WIDER THAN THE EIGHT:** the same path carries `UpdatePhaseLabel` (`updatePhaseLabels`, L63) and the HOLD sentence returned by `UpdateGuards.HoldFor`, including `backup.Manager.UpdateCopyHolds`'s *„csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem"* — **which is a PROMISE ABOUT WHETHER THE CUSTOMER'S FILES COME BACK**, and ranks this with R-590 rather than below it. **NOT closed by v0.260.0:** that release routed the 409 REFUSAL through `errText` (so the pre-flight refusals reach an English household in English), but a refusal is the path where nothing happened — these are the sentences for when something DID. **Fix shape, and the pattern already exists in this repo:** `UpdateError` stores a KEY plus args, exactly as v0.259.0's `degradedMessageFor` was changed to return a key, and the page resolves it with `errText`/`msg` at render — the decision stays language-free in one place while the words are chosen by whoever knows the reader. `util.MsgErrorf` already carries key+args across that gap. Render test per sentence in both languages. **Every one of these is BORN AS A KEY territory, so `i18n_go_keys.json` accounting applies.** | **READY — rank P2-MEDIUM; owner: CC (controller)** |
| **R-607** | **[P3-LOW] A forced catalog sync answers „nincs változás" while the cache DOES change, and `catalog_images` stays stale until a separate rescan — so the update badge can be wrong for a window nobody bounds.** MEASURED 2026-09-21 on scratch guest 9202 (controller v0.260.0) while proving R-524. A real catalog move was pushed, `POST /api/sync` was invoked, and it answered **„nincs változás"** — yet the box's own cache file `<data>/catalog-cache/templates/uptime-kuma/docker-compose.yml` **had moved to the new tag**. `Stack.CatalogImages` as served by the API stayed at the OLD value until a separate `POST /api/stacks/rescan`. **WHY IT MATTERS AND WHY IT IS NOT COSMETIC:** `CatalogImages` is the single input `stacks.CatalogOrder` compares against (v0.260.0), so for that window the badge answers from a stale catalog — it can read „Naprakész" on an app that IS behind, which is the exact failure §5.6 of `09` was written to prevent, arriving by a different door. **It also cost a measurement:** the session that found it nearly recorded a `tag-ok` badge as proof of the R-524 ahead arm when the badge was in fact stale; the honest reading came only after the rescan. **An instrument that can report an old value as a current one is not a measurement.** **TWO SEPARATE QUESTIONS, and the row does not conflate them:** (a) why the sync REPORTS no change when the working tree moved — a wrong sentence, possibly a comparison against the wrong ref; (b) whether `CatalogImages` is refreshed by the sync at all or only by `ScanStacks` on its own timer — if the latter, the staleness window is the scan interval and is bounded but unstated. **Neither was isolated** — this row records the observation, not a diagnosis. **Also observed in the same run, NOT diagnosed and folded in here rather than filed twice:** a removed app leaves `applied-compose.yml` behind in its stack directory. Stated as observed; it was not established whether that is intended. **Fix shape:** first reproduce with a loop that pushes a tag, syncs, and reads `catalog_images` on a timer, so the window is a NUMBER before anything is changed. Evidence: `audits/update-arc-2026-09-21/06-sync-after-bump.txt` and `16-r524-sync-box-ahead.txt`. | **READY — rank P3-LOW; owner: CC (controller)** |
| **R-608** | **[P2-MEDIUM] The controller swaps ITSELF in the middle of a guarded app update, and 04:30 sits inside the window proposed for automatic app updates.** FOUND 2026-09-21 by reading the clock, not by a failure. The controller self-updates daily at `self_update.auto_update_time` — **default 04:30** (`config/config.go` L422, scheduled `cmd/controller/main.go` ~L1365) — and again from `MaybeAutoUpdate` after ANY hub report once a floor sits above the box, so at any hour. The swap restarts the controller container. `09` §3b **Q1** proposes **02:30–05:00** for automatic app updates. **It contains 04:30.** **MEASURED, and the gap was NARROWER than it first looked — which is why the fix is where it is:** the updater's only busy gate was `backupRunning` (`updater.go` L61, read at L487 dry-run, L512 `TriggerUpdate`, L660 `maybeAutoUpdate`), wired in `main.go` L659 to `backupMgr.IsRunning()`. The guarded update's **`backing-up` phase DOES take the backup single-flight** (`RunAppBackupNow` → `acquireRunning`, `backup/update_guard.go:333`), so that ONE phase was already covered. `checking`, `safety-dump`, `pinning`, `pulling`, `starting` and `verifying` were not — and the last two are exactly where the new version may already have touched the customer's data. The reverse was absent too: `UpdatePreflight` never asked whether a swap was running. **CLOSED 2026-09-21 — controller v0.261.0.** `stacks.Manager.AnyUpdating()` → `Updater.SetAppUpdatingCheck`, a deliberate sibling of `SetBackupRunningCheck` consulted in the SAME three places; `Updater.IsUpdateRunning` → `Manager.SetSelfUpdatingCheck`, and `UpdatePreflight` refuses `self_updating`. Both wired in `main.go`, the only place holding both objects — **`stacks` never imports `selfupdate`.** Two sentences born as bundle keys. **THE PROPERTY THAT MATTERS MOST IS THAT THE LOCK DOES NOT LATCH:** `Stack.Updating` is cleared on done, failed AND held, so a HELD app does not block the controller's own updates — including the release that might fix whatever held it. A latching gate would be a worse failure than the one prevented, and a silent one. Pinned by `TestR608_LockReleasesAfterHold`. Four red-proofs, each seen to fail. | **CLOSED 2026-09-21 — controller v0.261.0** |
| **R-609** | **[P3-LOW] An update refusal has a machine-readable reason inside the process and none on the wire, so an unattended caller cannot tell "wait" from "never".** `UpdateRefusal.Reason` has existed since v0.237.0 (`busy`, `deploying`, `updating`, `held`, `migrating`, `memory`, `disk`, `no_backup`, `downgrade`, and `self_updating` since v0.261.0) and never left the process: the 409 body carried only the translated sentence. **The distinction is not decorative** — `busy`/`updating`/`deploying`/`migrating`/`self_updating` are TRANSIENT and `held`/`downgrade` are TERMINAL until a person acts. A caller that cannot tell them apart either gives up on a passing backup window or presses a terminally-refused button on every pass for ever. `09` §6.2's unattended caller reads exactly this. **CLOSED 2026-09-21 — controller v0.261.0.** The body gains `data: {"reason": "<Reason>"}`, ADDITIVELY; the sentence is unchanged and no page moves. Table-driven test over five reachable paths plus a control that a non-refusal carries none. **FOUND WHILE WRITING THE TEST, NOT BY READING — and it was the reason that matters most:** `actionStack` refuses a HELD app on **its own line, BEFORE `UpdatePreflight`** (`api/router.go` ~L601), so `held` would have been the one reason missing from the wire. That line now carries it too. Red-proof seen to fail. | **CLOSED 2026-09-21 — controller v0.261.0** |
| **R-610** | **[P2-MEDIUM] The power cut AFTER the new version has started was never measured — R-520 closed on the safe half only.** R-520 (CLOSED 2026-09-21) cut in `pulling`, where **nothing had run**: the pin goes back and that is the easy case. Its own last lines said the `starting` cut was "NOT measured and does not re-open this row", and no row carried it. The dangerous case is the cut AFTER the new version has started and may already have migrated the customer's data — where `RecoverUpdates` (`stacks/update.go:909`) marks the app Updating and RESUMES rather than rolling back. **That behaviour was READ from the source and never observed.** **CLOSED 2026-09-21 — measured THREE times on guest 9202, controller v0.260.0**, with three different apps and two different cut mechanisms: vikunja 2.3.0→2.6.0 and uptime-kuma 2.4.0→2.5.0 by `pct stop` (a real power cut), and wishlist v0.66.0→v0.67.0 by restarting ONLY the controller container (exactly what a self-update does). **All three ended HONEST:** the recovery line appeared, the update resumed, each app came up on the NEW version, and in every case all FOUR version observables agreed — `pinned_images`, `installed_images`, the live compose `image:` line, and `docker inspect` of the running container (digests matched too). No hold, no stuck `Updating`, no surviving journal, no retry loop. The seeded data read back through each app's own front door for the two where a post-cut read-back was taken. **THE DANGEROUS CASE WAS GENUINELY EXERCISED, and the proof is a log line, not an assumption:** vikunja's own log shows `Ran all migrations successfully` and `Vikunja version v2.6.0` at **12:28:26.881 UTC — 0.64 s after the cut decision and ~0.4 s before the guest stopped answering.** The 2.6.0 schema migration had ALREADY been applied to the customer's SQLite database when the power went. Recovery resumed FORWARD, so old-binary-on-migrated-database never happened — **but this branch is one step from it: had the cut landed a second earlier, in `pinning` or `pulling`, the 2.3.0 pin would have been put back onto a 2.6.0 database.** That is not a defect today; it is the reason §4's "no automatic rollback" ruling is right, and it is now evidence rather than argument. **INSTRUMENT LIMIT, stated because it bounds the claim:** `starting` lasts well under a second on this box. Three attempts across two cut mechanisms (`pct stop` returning in 3.0–3.8 s; a controller restart in 1.7 s) ALL landed in `verifying`. No phase was faked. **`RecoverUpdates` handles `starting` and `verifying` in ONE branch, so all three runs exercise the same recovery arm** — the arm under test. A cut that lands inside `starting` itself remains unmeasured and would need an in-process fault injector. Evidence: `audits/update-arc-gaps-2026-09-21/` 04, 05, 07. | **CLOSED 2026-09-21 — measured three times** |
| **R-611** | **[P3-LOW] A session reported "everything is done" over a phase it had silently skipped — the process failure, not the missing measurement.** The 2026-09-21 update-arc session's brief contained a Phase 5 spike: one app updated by the box with **nobody pressing anything**, once succeeding and once forced to fail. **It did not run, and nothing said so** — no evidence file, no code, and no sentence in `STATUS.md`, `UPDATE-ARC-STATE-2026-09-21.md`, either `REPORT.md` or `09`. `09` §6.2 was left describing Slice 6 "as it would be built" with no measurement under it, which reads like a considered design rather than an untested one. **Why this is a row and not a grumble:** the missing measurement was recoverable in an afternoon; the missing SENTENCE was not, because the next reader had no way to know it was missing. A skipped phase that is declared costs one line; a skipped phase that is not costs the next session its baseline. **CLOSED 2026-09-21 by the successor session**, which ran it (`audits/update-arc-gaps-2026-09-21/`, scenarios F and G) **and** adopted the standing habit that closes it generally: **the report's FIRST section is "not done", even when empty.** Every part and scenario of a brief is listed there if it was skipped, shortened or changed, with the reason. | **CLOSED 2026-09-21 — run by the successor session; "not done" is now the report's first section** |
| **R-612** | **[P1-HIGH] `wishlist` cannot be signed up to on a fresh Felhom install, the deploy reports SUCCESS, and the error the customer sees is a LIE.** MEASURED 2026-09-21 on guest 9202 while seeding for the power-cut drill. The image's first-boot `pnpm prisma db seed` is **`Killed` — OOM at the catalog's `mem_limit: 128M`**. Without it the `Role` and `Group` rows are absent, so **every** signup fails. **The message the user is shown is `User with username or email already exists`** while the container log says the real cause: `FOREIGN KEY constraint violated`. A household would conclude the account already exists and try to recover a password that was never created. **The controller reports the app running and HEALTHY throughout, and the deploy reported successful** — so nothing on the box says anything is wrong. Repaired on the scratch guest only, to unblock seeding: memory raised to 512 M, the image's own seed re-run, memory put back to 128 M. **The catalog was NOT changed** — the fix is a memory-limit question for the catalog and is deliberately left to a session that can measure the real ceiling rather than guess it. **Needs: the actual peak RSS of that seed, then a `mem_limit` that clears it, plus a check that the seed's failure is not silent.** | **READY — rank P1-HIGH; owner: CC (catalog + a look at whether a failed first-boot seed can ever be visible)** |
| **R-613** | **[P2-MEDIUM] `uptime-kuma` parks on its setup wizard with no login and no monitors, and the box tells the household it is HEALTHY.** MEASURED 2026-09-21 on guest 9202. On first boot uptime-kuma 2.4.0 sits at `SETUP-DATABASE` (`Waiting for user action...`) and its main socket.io server never starts. **The controller's `http :3001` probe sees the wizard's 302 and records the app as running and healthy.** So a monitoring app that cannot be logged into, and is monitoring nothing, is presented to the customer as fine. Passed through its own front door for the drill with `POST /setup-database {"dbConfig":{"type":"sqlite"}}`. **This is the health-check class the catalog skill already warns about — a probe that proves the PORT answers, not that the APP works** — and it is worth a row because the failure direction is a false GREEN, which no alarm will ever catch. **Needs: a healthcheck for this template that fails while the wizard is up** (the catalog `REUSE.md` maps the families), and a sweep for other templates whose probe would pass on a setup wizard. | **READY — rank P2-MEDIUM; owner: CC (catalog)** |
| **R-614** | **[P3-LOW] A stale `update_phase` survives a remove and redeploy, so a freshly installed app can read "Frissitve" before it has ever been updated.** OBSERVED 2026-09-21 on guest 9202: a newly deployed `uptime-kuma` read `update_phase=done` / „Frissitve" before any update had been run against it — left in the manager's IN-MEMORY stack state by the previous session's update of a since-removed instance of the same name. It cleared on the next controller restart. **Why it matters beyond the cosmetic:** `update_phase` is one of the fields a person (and, after R-609, an unattended caller) reads to decide whether an update happened. A value that outlives the app it described is the same class as the R-166 "absent means unknown" family — a confident answer about something that no longer exists. **Fix shape:** clear the update fields when a stack is removed, beside wherever `Updating`/`updateHeld` are reset; a test that removes and redeploys and asserts the phase is empty. Small. | **READY — rank P3-LOW; owner: CC (controller)** |
| **R-602** | **[P3-LOW] The language a signed-in page uses is NOT the language a cookie asks for, and a live probe that forgets this reports a fixed defect as unfixed.** FOUND 2026-09-21 verifying R-598 on guest 9201. `GET /backups` with `felhom_lang=en` returned the **Hungarian** page. That is correct — `langFor` step 2 says a request carrying a session reads the household's saved setting and deliberately ignores the visitor cookie, so a signed-in family never sees a language a previous visitor picked on the sign-in page of the same browser — but it means **the cookie is the right instrument for the anonymous claim/login/bind pages and the wrong one for every page behind auth**, where `?lang=` is. A session that had run only the cookie probe would have concluded R-598 was still open and fixed it a second time. **This is a documentation gap, not a code defect**, and it is the kind that costs a whole session: nothing in `10-localisation.md` §2.2 or in any runbook tells a prober which instrument to use where. **Fix shape:** four lines in `10-localisation.md` §2.2 — a table of surface → language instrument — and a pointer from the live-validation section of the workspace rules. Recorded meanwhile in `audits/i18n-closing-2026-09-21/live/backups-page.md`. | **READY — rank P3-LOW; owner: CC (docs)** |
| **R-603** | **[P3-LOW] An English string containing an apostrophe silently never matches on a rendered page, and a `strings.Contains` assertion reads exactly like a missing sentence.** FOUND 2026-09-21 while writing the R-598 render tests. `backup.target.absent` was first written as *"The system backup's drive cannot be reached…"*; `html/template` escapes `'` to `&#39;`, so the page carried the sentence and every assertion for it failed. **The failure mode is the expensive part:** the test said *"the English absent-drive copy never reached the page"*, which is indistinguishable from the handler not being wired — and the obvious next move is to go and re-fix the handler. Reworded to avoid the possessive, and all 23 new English values were then swept for `' " < > &` (zero). **The Hungarian bundle has never hit this** because Hungarian copy uses „quotes" and few apostrophes; **English copy will hit it again.** **Fix shape:** either a bundle gate that refuses an HTML-escapable character in a value destined for a page (and an allow-list for the ones that legitimately need one), or a test helper that compares against `html.EscapeString(want)` so the assertion cannot be fooled. The gate is the better shape — the helper only protects tests that remember to use it. | **READY — rank P3-LOW; owner: CC (controller)** |
| **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured, and a class-A app carries one sentence saying where its files ARE protected. Proven on demo-hp through the page the customer opens: Paperless-ngx reads „1. mentés … DB + Konfig" with „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi …", while its „2. mentés" row still reads „DB + Konfig + Adatok". Red-proof: restoring the old app-shaped label fails `TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold`. **RE-PROVEN 2026-09-16 on a FRESH box** (installed from the built ISO 1.28.0, controller 0.244.0, off-site on by default): the Nextcloud row read „1. mentés … DB + Konfig" with the new sentence, „2. mentés … Nincs 2. (off-drive) másolat", „3. mentés Sikeres restic → …your-storagebox.de"; „DB + Konfig + Adatok" appeared ZERO times while the local unit held no file leg. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** |