Undo bake-off: copy the folder wins (09 §3 decisions 19-20, §6.1a)
gates / gates (push) Successful in 26s

- 09 §3: decision 19 (the copy method is chosen by a bake-off) and 20 (the
  full-system backup waits for the update leg, inside its window; built later).
- Bake-off on 9202, docmost / romm / vikunja: both methods pass every case;
  the folder copy wins because an app with no database server gets no dump,
  so dump-and-load would need the folder copy anyway. 1-5 s extra downtime,
  ~420 MB/s, disk = the volumes.
- R-645 filed: lifting an update hold by hand lets the recovery unit be
  re-captured with the failed definition within seconds.

Documents and evidence only; product code follows in the controller.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-23 10:28:56 +02:00
parent 4c92beab8f
commit 5a349d9884
37 changed files with 2390 additions and 0 deletions
@@ -265,6 +265,21 @@ builds them (`audits/update-rulings-2026-09-23/`).
18. **Fleet view** (Q7). The report carries, per compose service (so the database too), the installed
reference, the catalog reference and the badge state. **Built later**, when the fleet grows.
### 2026-09-23 (afternoon) — two more operator rulings
19. **The undo's copy method is chosen by a bake-off, not by assumption.** Two methods are measured on
the same apps — *dump and load* (the database as a text file, loaded back) and *copy the folder*
(the app's named volumes copied with its containers stopped, put back on failure). The simpler one
that passes every case is built; if both pass, the folder copy wins on simplicity unless its
downtime or disk cost fails the bake-off's limits. **Why:** the morning spike found two traps in the
dump route (a cut-off file loads as success; a failed MariaDB load is half old, half new) and the
folder route had not been measured at all. Result: §6.1a.
20. **On update nights, the full-system backup waits for the update leg, inside its own window**
(answers R-643 / §6.4 part 7's open point). The leg stops starting new steps at **W+5h**, so the
full-system backup keeps at least one hour of its [W+2h, W+6h) window. **Built with §6.4 part 7,
not before.**
**RomM follow-ups, operator-agreed the same day:** the test bench watches memory after an update
(`upgrade-test.py`, 2026-09-23); a version move checks the memory limit (gate or checklist — §6.4);
R-636's louder repeated alarm.
@@ -798,6 +813,19 @@ product path that loads a safety dump back exists (`rollbackSafetyDump`), but on
restore calls it. The eight things the build must add are listed in the audit; §6.4 part 1 prices
them.
**THE COPY METHOD, chosen by the bake-off of 2026-09-23 afternoon (decision 19):
COPY THE FOLDER.** `audits/undo-bakeoff-2026-09-23/README.md`. Both methods passed every case on all
three apps (seeds before and after the backup, ledger equal, a cut-off copy caught before anything is
swapped or loaded). The folder copy wins because an app with no database server gets no dump at all,
so the dump route would have needed the folder copy anyway. **Shape:** after the pull and just before
`up` — where the app is stopped anyway to be recreated — each NAMED volume the app owns is copied with
`cp -a` into a sibling volume `<volume>.pre-update-<stamp>` by a helper container, which writes a
finished-marker last; a bind-mounted user folder is never copied and never touched. On a failed health
check the copy is validated (marker present) and put back, the old definition and pin are restored
from the journal's own copies, and the old version is checked with the OLD `.felhom.yml` probe.
Measured cost: ≈ 1–5 s extra downtime for the three apps, ≈ 420 MB/s, disk = the volumes' size (the
update refuses before moving anything when the copy would breach the 2 GB floor).
**The release could not reach the fleet by floor — R-472.** The hub holds a controller floor above the
vouched golden (publish-train rule 1), so under the weekly golden cadence (R-468) v0.237.0 and v0.238.0
were hand-deployed to the demo guests. **RESOLVED by §3 decision 7 (hub v0.112.0):** v0.239.0 reached
@@ -0,0 +1,10 @@
git:
branch: main
repo_url: https://gitea.dooplex.hu/admin/app-catalog-drill.git
sync_interval: 15m
token: <redacted>
username: "admin"
hub:
update:
health_timeout: 90s
@@ -0,0 +1,10 @@
sync 200 {'ok': True, 'data': {'ok': True, 'message': 'Sablonok naprakészek — nincs változás'}, 'message': 'Sablonok na
control 1 (9202 follows drill): a837c3a DRILL: FROM states for the undo bake-off
image: docmost/docmost:0.95.0
image: vikunja/vikunja:2.3.0
image: rommapp/romm:5.0.0
control 2 (9201 on live): cfcfe52 upgrade-test: watch memory after the readback (harness v2, R-635/R-462)
control 3 (live main): cfcfe527842865aa35a6d2ae9d361872e36afc9a refs/heads/main
@@ -0,0 +1,8 @@
10:22:59 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/nextcloud']
10:22:59 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['NEXTCLOUD_ADMIN_PASSWORD', 'HDD_PATH']
10:22:59 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
10:23:44 [1] deployed, controller state=running, pinned={'nextcloud': 'nextcloud:34.0.1-apache', 'nextcloud-db': 'mariadb:12.3', 'nextcloud-redis': 'redis:7-alpine'}
10:23:44 deployed: True
10:23:48 nextcloud: occ user:add :: The account "drilladbeb4" was created successfully Display name set to "drilladbeb4"
10:23:51 nextcloud: seeded user drilladbeb4
10:23:51 seeded: True
@@ -0,0 +1,5 @@
10:26:02 300 WebDAV uploads through the front door: {'201': 300}
10:26:03 PROPFIND lists 300 drill files (readback through the app); http 207
10:26:06 449
185374751 /v
@@ -0,0 +1,13 @@
10:26:49 == D's cost on the same database: mariadb-dump while running
dump 0.44s 1062572 B
== F on the same database: stop, copy, start
stop 1.58s
cp -a 185 MB: 0.81s
tar 185 MB: 0.53s
start 0.63s
== a 2 GiB synthetic volume (random bytes, not app data) for the extrapolation
page cache could NOT be dropped (LXC) - warm-ish numbers
cp -a 2 GiB: 4.87s
tar 2 GiB: 4.55s
/dev/loop1 69G 48G 18G 74% /var/lib/docker
@@ -0,0 +1,16 @@
10:27:13 [X] stop -> 200 {'ok': True, 'message': 'Stack nextcloud stop completed'}
10:27:18 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/nextcloud tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó
10:27:18 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead
10:27:46 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'nextcloud', 'volumes_removed': ['nextcloud_nextcloud_db_data', 'nextcloud_nextcloud_html', 'nextcloud_nextcloud_redis_data'],
10:27:53 [X] after remove: deployed=False leftovers='/opt/docker/stacks/nextcloud'
10:27:56 removed docmost_docmost_postgres_data.pre-undo
removed docmost_docmost_redis_data.pre-undo
removed docmost_docmost_storage.pre-undo
removed romm_romm_config.pre-undo
removed romm_romm_db_data.pre-undo
removed romm_romm_redis_data.pre-undo
removed vikunja_vikunja_data.pre-undo
removed vikunja_vikunja_db.pre-undo
nextcloud drive folder removed
0
@@ -0,0 +1,59 @@
# The undo bake-off, 2026-09-23 — copy the folder (F) against dump and load (D)
`09` §3 decision 19. Venue: scratch guest **9202**, controller v0.262.1, drill catalog (reset to live
`cfcfe5278428` first; `a837c3a7d8bd` FROM states), `update.health_timeout: 90s`. Same three apps and
edges as the morning spike (`../update-rulings-2026-09-23/`), each a real migrating edge made to fail
a deliberately wrong probe in the DRILL template only. Seed A before the backup, seed B after it,
both through the app's own front door. Driver: `bakeoff.py` (stages per app: prep → fcopy → break →
update → undoF → update → undoD → cutoff). Time-boxed at 2 h; took ~40 min.
**Method, stated so it is not over-read.** Both undos were done BY HAND in the order the product
would take them. The hold was lifted with the operator CLI + a controller restart (the only exit that
exists) and the product's own Start supplied the app's secrets.
## The table
| criterion | limit | **F** docmost / romm / vikunja | **D** docmost / romm / vikunja |
|---|---|---|---|
| seed A and seed B read back after the undo | both, every app | **yes·yes / yes·yes / yes·yes** | yes·yes / yes·yes / — (no safety dump exists: R-641) |
| tables / ledger equal to before the update | equal | **42 · 48 = / alembic 0095 = ¹ / 36 · 117 =** | 42 · 48 = / **24 base tables** · 0095 = / — |
| a cut-off copy detected before anything is swapped or loaded | detected | **docmost 4 of 4 cuts** (container exit 137, 42.9–68.4 of 69.2 MB copied, no finished-marker); **romm** (105.8 of 160.2 MB, no marker); vikunja **could not be cut** — 2.9 MB completes before a kill lands ² | docmost and romm: the half file lacks `-- PostgreSQL database dump complete` / `-- Dump completed` → refused; vikunja — |
| extra downtime before the update | ≤ 30 s | stop + copy: **≈ 2.4 s / ≈ 5.3 s / ≈ 1 s** (stop 0.52 / 3.79 / 0.18 s; copies 0.4–1.0 s per volume incl. the helper's start) | 0 — the dump runs against the running database |
| rate | stated | cp -a **185 MB (Nextcloud's MariaDB) in 0.81 s; 2 GiB in 4.87 s ≈ 420 MB/s**; tar 0.53 s / 4.55 s. Page cache could NOT be dropped in the guest, so these are warm-ish. **A 5 GB database ≈ 12 s at this rate**; on a cold cache or a spinning system disk expect 25–50 s | `mariadb-dump` of the same Nextcloud: 0.44 s, 1.06 MB |
| disk needed | stated; refuse near the 2 GB floor | the app's named volumes: **51 MB / 175 MB / 2.5 MB** (Nextcloud 185 MB) | the dump: 136 KB / 63 KB / — (Nextcloud 1 MB) |
| product code the build needs | estimated | **≈ 300 lines**, one mechanism for every app class | ≈ 450: the marker, two engine-specific empty-then-load paths, **and F's own machinery anyway** for apps with no database server |
¹ romm's base-table count was not captured on the F run (a quoting bug in the evidence script, fixed
before the D run, which read 24). F puts the whole datadir back byte-for-byte (copy bytes = source
bytes, file counts equal, `romm-2-fcopy.txt`), and the migration ledger read back equal.
² The marker logic is the same one proven on the other two; a copy that finishes cannot be cut.
## The choice — F, copy the folder (decision 19)
**Both methods pass every row.** Decision 19 then says the folder copy wins on simplicity unless its
downtime or disk cost fails the limits — and neither does: ≤ 5.3 s for these three, and the disk
cost is refused near the 2 GB floor by construction. **The decisive fact is the third column:** an
app with no database server gets no safety dump, so D would have to carry F's volume copy anyway.
F is one mechanism; D is F plus two engine loaders.
**Copy tool: `cp -a` into a sibling volume** (`<volume>.pre-update-<stamp>`), not `tar`: the two ran at
the same speed, and a sibling volume puts the undo back with the same `cp -a` and needs no space on
the app's own drive.
## Three things the bake-off found that the build must respect
1. **The undo must keep its own copy of the old definition.** Lifting the hold made the box re-capture
docmost's recovery unit ten seconds later — with the NEW definition — and a pin-back that read the
unit started the new version on the restored data, which migrated again (`docmost-40-undoF.txt`,
invalid run, kept as evidence; redo `docmost-41`). R-639, seen live.
2. **Killing `docker run` does not stop the copy** — the container runs on and finishes
(`docmost-60-cutoff.txt`, first block). The product must judge a copy by the helper container's own
exit status and the finished-marker it writes last, never by the client.
3. **Where the copy is taken decides the downtime.** Taken after the pull and just before `up` — where
the update stops the app anyway to recreate it — the extra downtime is the copy alone.
## Evidence
`docmost-*`, `romm-*`, `vikunja-*` (per stage), `70`–`73` (the rate test on Nextcloud, removed after),
`01`/`02` (repoint + three controls). Nextcloud and every `.pre-undo`/`.cut`/`.rate` volume were
removed by name at the end of the phase (`73-rate-teardown.txt`).
@@ -0,0 +1,266 @@
#!/usr/bin/env python3
"""bakeoff.py — `09` §3 decision 19: the two ways of keeping the undo's last-second copy, measured on
the same three apps (docmost / PostgreSQL, romm / MariaDB, vikunja / SQLite in a volume), guest 9202,
drill catalog.
F — copy the folder: the app's NAMED volumes copied with its containers stopped; put back on failure.
D — dump and load, as fixed by the morning spike: completion marker first; PostgreSQL drops and
recreates the dump's schemas inside the load's transaction; MariaDB drops every table, then loads.
EVIDENCE, NOT PRODUCT. Product acts go through the endpoints the UI invokes (deploy, backup, sync,
rescan, update, start, remove). The undo has no product path yet, so it is done by hand — the hold is
lifted with the operator CLI + a controller restart (the only exit that exists), and the product's own
Start supplies the app's secrets (they are never decrypted here).
Usage: python3 bakeoff.py <stage> <app> stages: prep fcopy break update undoF undoD cutoff state
"""
import json, os, sys, time
sys.path.insert(0, ".")
import walk as w
from spike import FX, SUB, load, save, ts, break_edge, docmost_seed_b, docmost_verify_b, romm_seed_b, \
romm_verify_b, vik_seed_b, vik_verify_b
EDGE = {"docmost": ("docmost/docmost:0.95.0", "docmost/docmost:0.96.0", 3000, 3999),
"romm": ("rommapp/romm:5.0.0", "rommapp/romm:5.3.0", 8080, 8999),
"vikunja": ("vikunja/vikunja:2.3.0", "vikunja/vikunja:2.6.0", 3456, 3999)}
UNIT = {"docmost": "/mnt/sys_drive/felhom-data/backups/primary/docmost",
"vikunja": "/mnt/sys_drive/felhom-data/backups/primary/vikunja",
"romm": "/mnt/felhom-drives/scratch_hdd/userdata/romm/backups/primary/romm"}
SUFFIX = ".pre-undo"
def seed_b(app, sub, A):
return {"docmost": docmost_seed_b, "romm": romm_seed_b, "vikunja": vik_seed_b}[app](sub, A)
def verify_b(app, sub, A, B):
r = {"docmost": docmost_verify_b, "romm": romm_verify_b, "vikunja": vik_verify_b}[app](sub, A, B)
return all(r.values()) if isinstance(r, dict) else r
def db_state(app):
"""The table count and the migration ledger, asked of the engine (or the SQLite file, read-only)."""
if app == "docmost":
return w.guest("docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1 | tr '\\n' ' '; "
"docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*)||' ledger, newest '||max(name) from kysely_migration\" 2>&1").strip()
if app == "romm":
q = lambda sql: f"docker exec romm-db sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" romm' 2>&1 | tr '\\n' ' '"
base_tables = q("select count(*) from information_schema.tables where table_schema=database() and table_type=0x42415345205441424C45")
alembic = q("select version_num from alembic_version")
return w.guest(f'echo "tables=$({base_tables}) alembic=$({alembic})"').strip()
return w.guest("""python3 - <<'PY'
import sqlite3
c=sqlite3.connect('file:/var/lib/docker/volumes/vikunja_vikunja_db/_data/vikunja.db?mode=ro',uri=True)
t=c.execute("select count(*) from sqlite_master where type='table'").fetchone()[0]
m=c.execute("select count(*), max(id) from migration").fetchone()
print(f"tables={t} ledger={m[0]} newest={m[1]}")
PY""").strip()
def volumes(app):
return [v for v in w.guest(f"docker volume ls -q --filter label=com.docker.compose.project={app}").split()
if not v.endswith(SUFFIX)]
def containers(app):
return w.guest(f"docker ps -a --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'").split()
def stage_prep(app):
s = {"app": app, "sub": SUB[app]}
w.say(f"=== prep {app}: deploy at the drill FROM pin")
w.say("deployed:", w.deploy(app, s["sub"]))
s["seedA"] = FX[app].seed(w, s["sub"], w.say)
w.say("C1 A:", FX[app].verify(w, s["sub"], s["seedA"], w.say))
w.backup_now(app)
w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60)
s["seedB"] = seed_b(app, s["sub"], s["seedA"])
w.say("B reads back:", verify_b(app, s["sub"], s["seedA"], s["seedB"]))
s["volumes"] = volumes(app)
w.say("named volumes:", s["volumes"])
save(app, s)
def stage_fcopy(app):
"""F at safety-dump time: stop, copy every named volume (cp -a into a sibling volume, and
separately a tar, for the rate), start. Measures the EXTRA downtime and the disk."""
s = load(app)
s["state_before_update"] = db_state(app)
w.say("db state before the update:", s["state_before_update"])
vols = s["volumes"]
w.say(w.guest(f"cp /opt/docker/stacks/{app}/docker-compose.yml /root/pre-{app}-compose.yml; grep -m1 'image: {EDGE[app][0]}' /root/pre-{app}-compose.yml"))
script = f"""
set -u
cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}')
t0=$(date +%s.%N)
docker stop $cs >/dev/null
t1=$(date +%s.%N)
for v in {' '.join(vols)}; do
docker volume rm -f "$v{SUFFIX}" >/dev/null 2>&1
docker volume create "$v{SUFFIX}" >/dev/null
a=$(date +%s.%N)
docker run --rm -v "$v":/from:ro -v "$v{SUFFIX}":/to alpine:3.20 sh -c 'cp -a /from/. /to/ && sync' || echo "COPY FAILED $v"
b=$(date +%s.%N)
bytes=$(docker run --rm -v "$v":/from:ro alpine:3.20 du -sb /from | cut -f1)
files=$(docker run --rm -v "$v":/from:ro alpine:3.20 sh -c 'find /from | wc -l')
cbytes=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 du -sb /to | cut -f1)
cfiles=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 sh -c 'find /to | wc -l')
mkdir -p /root/bk; c=$(date +%s.%N)
docker run --rm -v "$v":/vol:ro -v /root/bk:/out alpine:3.20 tar cf /out/$v.tar -C /vol . ; d=$(date +%s.%N)
tb=$(stat -c %s /root/bk/$v.tar); rm -f /root/bk/$v.tar
python3 -c "print('VOL $v bytes=$bytes files=$files | copy bytes=$cbytes files=$cfiles | cp -a %.2fs | tar %.2fs (%s B)' % ($b-$a, $d-$c, '$tb'))"
done
t2=$(date +%s.%N)
docker start $cs >/dev/null
t3=$(date +%s.%N)
python3 -c "print('stop %.2fs copy(all, cp -a + the tar measurement) %.2fs start %.2fs' % ($t1-$t0, $t2-$t1, $t3-$t2))"
df -B1 --output=avail /var/lib/docker | tail -1 | awk '{{printf "free on the docker root: %.2f GiB\\n", $1/2^30}}'
"""
out = w.guest(script, timeout=1800)
w.say(out)
t0 = time.time()
w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2)
w.say(f"front door answering again {round(time.time()-t0,1)}s after the start")
s["fcopy"] = out
s["fcopy_at"] = ts()
save(app, s)
def stage_break(app):
s = load(app)
frm, to, p0, p1 = EDGE[app]
s["break_commit"] = break_edge(app, frm, to, p0, p1)
w.say("badge caught up after", w.sync_rescan(app, to), "s")
save(app, s)
def stage_update(app):
s = load(app)
r = w.press_update(app, poll=1.0)
w.say(json.dumps({k: r[k] for k in ("final_phase", "hold_reason", "state", "duration_s")}, ensure_ascii=False))
s.setdefault("updates", []).append(r)
save(app, s)
LIFT = """docker exec felhom-controller /usr/local/bin/felhom-controller --clear-restore-hold {app} 2>&1 | grep -E 'CLEARED|no restore hold'
docker restart felhom-controller >/dev/null; sleep 15"""
# The OLD definition comes from the bake-off's OWN pre-update copy (/root/pre-<app>-compose.yml, taken
# in fcopy), NEVER from the recovery unit: measured 2026-09-23 10:59, the unit was re-captured 10 s
# after the hold was lifted — with the NEW definition — and a pin-back that read it started the new
# version again. That is R-639 seen live, and why the product undo keeps its own copies.
PINBACK = """S=/opt/docker/stacks/{app}; P=/root/pre-{app}-compose.yml
grep -m1 'image: .*{frm}' $P >/dev/null || {{ echo "PRE-UPDATE COPY MISSING OR WRONG: $P"; exit 1; }}
cp $P $S/applied-compose.yml; cp $P $S/docker-compose.yml
sed -i 's#^\\(\\s*{svc}: \\){to}$#\\1{frm}#' $S/app.yaml; sed -n '/^pinned_images:/,$p' $S/app.yaml | head -4"""
def lift_and_pinback(app):
frm, to, _, _ = EDGE[app]
t = time.time()
w.say(w.guest(LIFT.format(app=app)))
w.say(w.guest(PINBACK.format(unit=UNIT[app], app=app, svc=app, frm=frm, to=to)))
w.login()
return round(time.time() - t, 1)
def stage_undoF(app):
s = load(app)
w.say(f"=== undo by F: {app}")
w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)")
vols = s["volumes"]
script = f"""
t0=$(date +%s.%N)
cs=$(docker ps -a --filter label=com.docker.compose.project={app} -q); [ -n "$cs" ] && docker stop $cs >/dev/null
for v in {' '.join(vols)}; do
docker run --rm -v "$v{SUFFIX}":/from:ro -v "$v":/to alpine:3.20 sh -c 'rm -rf /to/..?* /to/.[!.]* /to/* ; cp -a /from/. /to/ && sync' || echo "RESTORE FAILED $v"
done
python3 -c "import time;print('volumes put back in %.2fs' % (time.time()-$t0))"
"""
w.say(w.guest(script, timeout=1800))
t0 = time.time()
c, d = w.ctl("POST", f"/api/stacks/{app}/start")
w.say("product start ->", c)
up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2)
th = round(time.time() - t0, 1)
A = FX[app].verify(w, s["sub"], s["seedA"], w.say)
B = verify_b(app, s["sub"], s["seedA"], s["seedB"])
st = db_state(app)
w.say(f"F RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})")
s["undoF"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st}
save(app, s)
def stage_undoD(app):
"""D, on the SECOND failed update: its safety dump (written by the product, phase 3) is the copy."""
s = load(app)
w.say(f"=== undo by D: {app}")
w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)")
if app == "vikunja":
w.say("vikunja has no database server: the product wrote NO safety dump (R-641). D's copy for it "
"is a volume tar at safety-dump time, which is F with extra steps — recorded, not re-run.")
return
c, d = w.ctl("POST", f"/api/stacks/{app}/start") # the product supplies the env; the app half will refuse
time.sleep(8)
if app == "docmost":
load_cmd = """{ echo 'DROP SCHEMA public CASCADE; CREATE SCHEMA public;'; cat "$D"; } | docker exec -i docmost-postgres psql -v ON_ERROR_STOP=1 --single-transaction -U docmost -d docmost > /root/dload.out 2>&1"""
marker = "-- PostgreSQL database dump complete"
else:
load_cmd = """{ echo 'SET FOREIGN_KEY_CHECKS=0;'; docker exec romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" -N -e "select concat(\\"DROP TABLE IF EXISTS \\`\\",table_name,\\"\\`;\\") from information_schema.tables where table_schema=\\"romm\\" and table_type=\\"BASE TABLE\\""' ; cat "$D"; } | docker exec -i romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" romm' > /root/dload.out 2>&1"""
marker = "-- Dump completed"
script = f"""
D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql | head -1); echo "undo copy: $(basename $D) $(stat -c %s $D) B"
docker stop {app} >/dev/null 2>&1; docker update --restart=no {app} >/dev/null
t0=$(date +%s.%N)
if tail -n 5 "$D" | grep -q -- '{marker}'; then echo "marker present"; else echo "MARKER ABSENT - refusing"; exit 0; fi
{load_cmd}; echo "load rc=$?"; tail -2 /root/dload.out
python3 -c "import time;print('marker check + load %.2fs' % (time.time()-$t0))"
docker update --restart=unless-stopped {app} >/dev/null; docker start {app} >/dev/null
"""
w.say(w.guest(script, timeout=900))
t0 = time.time()
up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat"}[app], want=("200",), tries=60, delay=2)
th = round(time.time() - t0, 1)
A = FX[app].verify(w, s["sub"], s["seedA"], w.say)
B = verify_b(app, s["sub"], s["seedA"], s["seedB"])
st = db_state(app)
w.say(f"D RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})")
s["undoD"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st}
save(app, s)
def stage_cutoff(app):
"""The cut-off copy, both methods, DETECTED before anything is loaded or swapped.
F: a copy killed half-way (timeout) — compared with the source and checked for the finished-marker
the build would write only after cp exits 0. D: the safety dump cut in half — the marker check."""
s = load(app)
big = max(s["volumes"], key=lambda v: int(w.guest(f"docker run --rm -v {v}:/v:ro alpine:3.20 du -sb /v | cut -f1").strip() or 0))
script = f"""
v={big}
docker volume rm -f $v.cut >/dev/null 2>&1; docker volume create $v.cut >/dev/null
cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'); docker stop $cs >/dev/null
# Kill the copy CONTAINER, not the client (killing `docker run` leaves the container copying — measured
# on docmost 2026-09-23, rc 137 with a complete copy and a marker).
docker run -d --name cutcopy -v $v:/from:ro -v $v.cut:/to alpine:3.20 sh -c 'cp -a /from/. /to/ && touch /to/.felhom-copy-complete' >/dev/null
sleep 0.02; docker kill cutcopy >/dev/null 2>&1; echo "copy container exit=$(docker wait cutcopy) (137 = killed)"; docker rm -f cutcopy >/dev/null 2>&1
echo "source bytes=$(docker run --rm -v $v:/f:ro alpine:3.20 du -sb /f | cut -f1) cut copy bytes=$(docker run --rm -v $v.cut:/f:ro alpine:3.20 du -sb /f | cut -f1)"
echo "finished-marker in the cut copy: $(docker run --rm -v $v.cut:/f:ro alpine:3.20 sh -c 'ls /f/.felhom-copy-complete 2>/dev/null | wc -l') -> F refuses to swap"
docker volume rm -f $v.cut >/dev/null; docker start $cs >/dev/null
D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql 2>/dev/null | head -1)
if [ -n "$D" ]; then head -c $(( $(stat -c %s $D) / 2 )) $D > /root/cut.sql
echo "D: whole copy marker: $(tail -n5 $D | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') cut copy marker: $(tail -n5 /root/cut.sql | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') -> D refuses to load"
rm -f /root/cut.sql
else echo "D: no safety dump exists for this app (R-641)"; fi
"""
w.say(f"=== cut-off copy: {app} (largest volume {big})")
out = w.guest(script, timeout=600)
w.say(out)
s["cutoff"] = out
save(app, s)
if __name__ == "__main__":
w.login()
{"prep": stage_prep, "fcopy": stage_fcopy, "break": stage_break, "update": stage_update,
"undoF": stage_undoF, "undoD": stage_undoD, "cutoff": stage_cutoff,
"state": lambda a: w.say(db_state(a))}[sys.argv[1]](sys.argv[2])
@@ -0,0 +1,13 @@
09:53:18 === prep docmost: deploy at the drill FROM pin
09:53:18 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
09:54:58 [1] deployed, controller state=running, pinned={'docmost': 'docmost/docmost:0.95.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'}
09:54:58 deployed: True
09:54:59 docmost: /api/auth/setup http=200 rc=0
09:54:59 docmost: login as the seeded user http=200 ok=True
09:54:59 C1 A: True
09:54:59 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'}
09:55:40 [4] backup idle; last=None
09:55:40 docmost seed B: /api/spaces/create http=200 {"data":{"id":"01a0cd43-623a-7ddf-9b62-213653bc5558","name":"drillB3aab00","description":"","slug":"drillb3aab00","logo":null,"visibility":"private","defaultRol
09:55:40 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False)
09:55:40 B reads back: True
09:55:43 named volumes: ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage']
@@ -0,0 +1,8 @@
09:55:53 db state before the update: 42 48 ledger, newest 20260620T010047-personal-spaces
09:56:05 VOL docmost_docmost_postgres_data bytes=51156449 files=1561 | copy bytes=51156449 files=1561 | cp -a 0.98s | tar 0.54s (51838464 B)
VOL docmost_docmost_redis_data bytes=39819 files=6 | copy bytes=39819 files=6 | cp -a 0.46s | tar 0.45s (37376 B)
VOL docmost_docmost_storage bytes=4096 files=1 | copy bytes=4096 files=1 | cp -a 0.46s | tar 0.39s (1536 B)
stop 0.52s copy(all, cp -a + the tar measurement) 9.00s start 0.75s
free on the docker root: 21.04 GiB
09:56:17 front door answering again 12.2s after the start
@@ -0,0 +1,9 @@
09:56:25 [5] drill commit 386a34da911a: docmost docmost/docmost:0.95.0 -> docmost/docmost:0.96.0 (push rc=0)
09:56:29 badge caught up after 4.5 s
09:56:30 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
09:56:30 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None
09:56:31 + 1.0s phase=pulling label=Új verzió letöltése… err=None hold=None
09:57:44 + 74.9s phase=starting label=Indítás az új verzióval… err=None hold=None
09:57:46 + 76.9s phase=verifying label=Működés ellenőrzése… err=None hold=None
09:59:18 + 168.1s phase=failed label=A frissítés nem sikerült err=A(z) docmost frissítése 2026-09-23 09:59-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 09:55 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza. hold=A(z) docmost frissítése 2026-09-23 09:59-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 09:55 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.
09:59:18 {"final_phase": "failed", "hold_reason": "A(z) docmost frissítése 2026-09-23 09:59-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 09:55 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.", "state": "stopped", "duration_s": 168.1}
@@ -0,0 +1,15 @@
09:59:23 === undo by F: docmost
09:59:41 [INFO] [settings] restore hold CLEARED for docmost
09:59:43 pinned_images:
docmost: docmost/docmost:0.95.0
docmost-postgres: postgres:16-alpine
docmost-redis: redis:7-alpine
09:59:43 hold lift + pin back took 20.2 s (not part of a product undo)
09:59:48 volumes put back in 2.20s
09:59:59 product start -> 200
10:00:12 docmost: login as the seeded user http=200 ok=True
10:00:12 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False)
10:00:15 F RESULT: healthy=True after 23.6s A=True B=True db now: 48 52 ledger, newest 20260904T171920-public-spaces (before the update: 42 48 ledger, newest 20260620T010047-personal-spaces)
@@ -0,0 +1,17 @@
10:00:57 === REDO of the F undo with the bake-off's OWN pre-update definition (the first try read the re-captured unit — see docmost-40)
10:01:00 image: docmost/docmost:0.95.0
10:01:02 pinned_images:
docmost: docmost/docmost:0.95.0
docmost-postgres: postgres:16-alpine
docmost-redis: redis:7-alpine
10:01:07 stop + volumes put back in 2.91s
10:01:18 product start -> 200
10:01:33 docmost/docmost:0.95.0 Up 15 seconds (healthy)
successfully started
10:01:34 docmost: login as the seeded user http=200 ok=True
10:01:34 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False)
10:01:36 F RESULT (redo): healthy=True after 23.6s A=True B=True db now: 42 48 ledger, newest 20260620T010047-personal-spaces (before the update: 42 48 ledger, newest 20260620T010047-personal-spaces)
@@ -0,0 +1,26 @@
10:01:43 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
10:01:43 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None
10:01:44 + 1.1s phase=pulling label=Új verzió letöltése… err=None hold=None
10:01:45 + 2.1s phase=starting label=Indítás az új verzióval… err=None hold=None
10:01:46 + 3.1s phase=verifying label=Működés ellenőrzése… err=None hold=None
10:03:18 + 95.4s phase=failed label=A frissítés nem sikerült err=A(z) docmost frissítése 2026-09-23 10:03-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 09:59 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza. hold=A(z) docmost frissítése 2026-09-23 10:03-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 09:59 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.
10:03:18 {"final_phase": "failed", "hold_reason": "A(z) docmost frissítése 2026-09-23 10:03-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 09:59 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.", "state": "stopped", "duration_s": 95.4}
10:03:18 === undo by D: docmost
10:03:36 [INFO] [settings] restore hold CLEARED for docmost
10:03:38 pinned_images:
docmost: docmost/docmost:0.95.0
docmost-postgres: postgres:16-alpine
docmost-redis: redis:7-alpine
10:03:38 hold lift + pin back took 20.2 s (not part of a product undo)
10:04:01 undo copy: pre-restore-20260923T080143Z-docmost-postgres.sql 136043 B
marker present
load rc=0
ALTER TABLE
ALTER TABLE
marker check + load 1.87s
10:04:13 docmost: login as the seeded user http=200 ok=True
10:04:14 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False)
10:04:16 D RESULT: healthy=True after 12.2s A=True B=True db now: 42 48 ledger, newest 20260620T010047-personal-spaces (before the update: 42 48 ledger, newest 20260620T010047-personal-spaces)
@@ -0,0 +1,16 @@
10:04:32 === cut-off copy: docmost (largest volume docmost_docmost_postgres_data)
10:04:39 copy rc=137 (137 = killed)
source bytes=69334709 cut copy bytes=69334709
finished-marker in the cut copy: 1 -> F refuses to swap
D: whole copy marker: 1 cut copy marker: 0 -> D refuses to load
10:04:55 NOTE: the attempt above killed the docker CLIENT (rc 137) while the copy CONTAINER ran on and finished — so nothing was cut and the marker was written. Redo: kill the copy CONTAINER itself mid-copy.
10:05:07 kill after 0.15s: container exit=0 source bytes=69172397 cut copy bytes=69172397 finished-marker=1
kill after 0.25s: container exit=0 source bytes=69172397 cut copy bytes=69172397 finished-marker=1
kill after 0.35s: container exit=0 source bytes=69172397 cut copy bytes=69172397 finished-marker=1
10:05:31 kill after 0s: container exit=137 source bytes=69172397 cut copy bytes=42921516 finished-marker=0
kill after 0.02s: container exit=137 source bytes=69172397 cut copy bytes=48586287 finished-marker=0
kill after 0.05s: container exit=137 source bytes=69172397 cut copy bytes=59159114 finished-marker=0
kill after 0.08s: container exit=137 source bytes=69172397 cut copy bytes=68445276 finished-marker=0
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,60 @@
#!/usr/bin/env python3
"""Point guest 9202 at the drill catalog (and a 90 s health timeout), or restore the saved config.
`09` §6.5: `git.repo_url` alone is INERT (R-615) — the cache dir must go too. The saved copy is
`controller.yaml.pre-bakeoff` (NOT the older `.pre-28`, which a restore must never pick up).
"""
import re, sys, io
sys.path.insert(0, '.')
import walk as w
VOL = "/var/lib/docker/volumes/felhom-controller-data/_data"
DRILL_REPO = "https://gitea.dooplex.hu/admin/app-catalog-drill.git"
def creds():
for l in io.open("/home/kisfenyo/.git-credentials").read().strip().split("\n"):
m = re.match(r'https://(admin):([^@]+)@gitea\.dooplex\.hu', l)
if m:
return m.group(1), m.group(2)
raise SystemExit("no admin credential")
def to_drill():
u, t = creds()
print(w.guest(f"""
set -e
test -f {VOL}/controller.yaml.pre-bakeoff || cp -p {VOL}/controller.yaml {VOL}/controller.yaml.pre-bakeoff
python3 - <<'PY'
import re
p = "{VOL}/controller.yaml"
s = open(p).read()
s = re.sub(r'(^\\s+repo_url: ).*$', r'\\g<1>{DRILL_REPO}', s, count=1, flags=re.M)
s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+token: ).*$', r'\\g<1>"{t}"', s, count=1, flags=re.M)
s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+username: ).*$', r'\\g<1>"{u}"', s, count=1, flags=re.M)
if not re.search(r'^update:', s, re.M):
s += "update:\\n health_timeout: 90s\\n"
open(p, "w").write(s)
PY
rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache
docker restart felhom-controller >/dev/null
sleep 15
grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: <redacted>/'
grep -A2 '^update:' {VOL}/controller.yaml
"""))
def restore():
print(w.guest(f"""
set -e
cp -p {VOL}/controller.yaml.pre-bakeoff {VOL}/controller.yaml
rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache
docker restart felhom-controller >/dev/null
sleep 15
grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: <redacted>/'
grep -c '^update:' {VOL}/controller.yaml || true
"""))
if __name__ == "__main__":
to_drill() if sys.argv[1] == "drill" else restore()
@@ -0,0 +1,15 @@
10:05:44 === prep romm: deploy at the drill FROM pin
10:05:47 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/romm']
10:05:47 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['HDD_PATH']
10:05:47 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
10:07:02 [1] deployed, controller state=running, pinned={'romm': 'rommapp/romm:5.0.0', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'}
10:07:02 deployed: True
10:07:12 romm: POST /api/users http=201
10:07:13 romm: login as the seeded user http=200 ok=True
10:07:13 C1 A: True
10:07:13 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'}
10:08:24 [4] backup idle; last=None
10:08:54 romm seed B: POST /api/users (as the admin) http=201 {"id":2,"username":"drillb12740b","email":"drillb12740b@gate.invalid","enabled":true,"role":"user","permission_group_id"
10:08:54 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False)
10:08:54 B reads back: True
10:08:57 named volumes: ['romm_romm_config', 'romm_romm_db_data', 'romm_romm_redis_data']
@@ -0,0 +1,10 @@
10:08:59 db state before the update: tables=TABLE" romm: 1: Syntax error: Unterminated quoted string alembic=0095_virtual_collections_source
10:09:01 image: rommapp/romm:5.0.0
10:09:15 VOL romm_romm_config bytes=4158 files=2 | copy bytes=4158 files=2 | cp -a 0.40s | tar 0.38s (2560 B)
VOL romm_romm_db_data bytes=160302192 files=273 | copy bytes=160302192 files=273 | cp -a 0.68s | tar 0.50s (160458240 B)
VOL romm_romm_redis_data bytes=14512195 files=6 | copy bytes=14512195 files=6 | cp -a 0.42s | tar 0.40s (14509056 B)
stop 3.79s copy(all, cp -a + the tar measurement) 7.75s start 0.68s
free on the docker root: 19.04 GiB
10:09:49 front door answering again 33.8s after the start
@@ -0,0 +1,2 @@
10:09:50 [5] drill commit de819653d3d0: romm rommapp/romm:5.0.0 -> rommapp/romm:5.3.0 (push rc=0)
10:09:55 badge caught up after 4.4 s
@@ -0,0 +1,7 @@
10:09:55 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
10:09:55 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None
10:09:56 + 1.0s phase=pulling label=Új verzió letöltése… err=None hold=None
10:10:08 + 13.3s phase=starting label=Indítás az új verzióval… err=None hold=None
10:10:14 + 19.5s phase=verifying label=Működés ellenőrzése… err=None hold=None
10:11:54 + 118.9s phase=failed label=A frissítés nem sikerült err=A(z) romm frissítése 2026-09-23 10:11-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:08 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem. hold=A(z) romm frissítése 2026-09-23 10:11-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:08 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.
10:11:54 {"final_phase": "failed", "hold_reason": "A(z) romm frissítése 2026-09-23 10:11-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:08 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.", "state": "stopped", "duration_s": 119.0}
@@ -0,0 +1,15 @@
10:11:54 === undo by F: romm
10:12:12 [INFO] [settings] restore hold CLEARED for romm
10:12:14 pinned_images:
romm: rommapp/romm:5.0.0
romm-db: mariadb:11.4
romm-redis: redis:7-alpine
10:12:14 hold lift + pin back took 20.2 s (not part of a product undo)
10:12:18 volumes put back in 1.64s
10:12:30 product start -> 200
10:13:05 romm: login as the seeded user http=200 ok=True
10:13:05 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False)
10:13:08 F RESULT: healthy=True after 45.3s A=True B=True db now: tables=TABLE" romm: 1: Syntax error: Unterminated quoted string alembic=0095_virtual_collections_source (before the update: tables=TABLE" romm: 1: Syntax error: Unterminated quoted string alembic=0095_virtual_collections_source)
@@ -0,0 +1,7 @@
10:13:09 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
10:13:09 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None
10:13:10 + 1.1s phase=pulling label=Új verzió letöltése… err=None hold=None
10:13:11 + 2.1s phase=starting label=Indítás az új verzióval… err=None hold=None
10:13:14 + 5.2s phase=verifying label=Működés ellenőrzése… err=None hold=None
10:14:48 + 99.5s phase=failed label=A frissítés nem sikerült err=A(z) romm frissítése 2026-09-23 10:14-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:11 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem. hold=A(z) romm frissítése 2026-09-23 10:14-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:11 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.
10:14:48 {"final_phase": "failed", "hold_reason": "A(z) romm frissítése 2026-09-23 10:14-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:11 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.", "state": "stopped", "duration_s": 99.5}
@@ -0,0 +1,17 @@
10:14:48 === undo by D: romm
10:15:06 [INFO] [settings] restore hold CLEARED for romm
10:15:08 pinned_images:
romm: rommapp/romm:5.0.0
romm-db: mariadb:11.4
romm-redis: redis:7-alpine
10:15:08 hold lift + pin back took 20.1 s (not part of a product undo)
10:15:40 undo copy: pre-restore-20260923T081309Z-romm-mariadb.sql 62943 B
marker present
load rc=0
marker check + load 1.41s
10:16:15 romm: login as the seeded user http=200 ok=True
10:16:16 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False)
10:16:18 D RESULT: healthy=True after 33.6s A=True B=True db now: tables=24 alembic=0095_virtual_collections_source (before the update: tables=TABLE" romm: 1: Syntax error: Unterminated quoted string alembic=0095_virtual_collections_source)
@@ -0,0 +1,6 @@
10:16:26 === cut-off copy: romm (largest volume romm_romm_db_data)
10:16:35 copy container exit=137 (137 = killed)
source bytes=160215912 cut copy bytes=105844788
finished-marker in the cut copy: 0 -> F refuses to swap
D: whole copy marker: 1 cut copy marker: 0 -> D refuses to load
@@ -0,0 +1,160 @@
#!/usr/bin/env python3
"""spike.py — Part 1 of the 2026-09-23 brief: the AUTOMATIC UNDO, performed BY HAND on guest 9202.
EVIDENCE, NOT PRODUCT. Every product act goes through the endpoints the UI invokes (deploy, backup,
sync, rescan, update, remove). The UNDO itself has no product path yet (that is what is being
spiked), so it is performed by hand with plain docker/compose inside the guest, using exactly the
steps the product would take, each one timed.
State between stages lives in state-<app>.json so each stage can be run, read, and only then
followed by the next (the hand undo needs a person looking at what the previous step left).
"""
import json, os, sys, time, re
sys.path.insert(0, ".")
import walk as w
import fixtures as fx
HERE = os.path.dirname(os.path.abspath(__file__))
FX = {"docmost": fx.Docmost(), "vikunja": fx.Vikunja(), "romm": fx.Romm()}
SUB = {"docmost": "docs", "vikunja": "tasks", "romm": "arcade"}
def st_path(app):
return os.path.join(HERE, f"state-{app}.json")
def load(app):
return json.load(open(st_path(app))) if os.path.exists(st_path(app)) else {}
def save(app, s):
json.dump(s, open(st_path(app), "w"), indent=2, ensure_ascii=False)
def ts():
return time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
# ---- a SECOND seed, written AFTER the backup and BEFORE the update. It is the discriminator: only
# the pre-pin safety dump can hold it — the backup tier copy was taken before it existed. So if it
# reads back after the undo, the undo used the safety dump; if only A reads back, it used the tier.
def docmost_seed_b(sub, A):
jar = "/tmp/dm.jar"
w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json",
data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST")
name = "drillB" + os.urandom(3).hex()
rc, code, out = w.app_curl(sub, "/api/spaces/create", "-b", jar, "-H", "Content-Type: application/json",
data=json.dumps({"name": name, "slug": name.lower()}), method="POST")
w.say(f" docmost seed B: /api/spaces/create http={code} {out[:160]}")
return {"space": name} if code in ("200", "201") else None
def docmost_verify_b(sub, A, B):
jar = "/tmp/dm.jar"
rc, code, out = w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json",
data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST")
if code not in ("200", "201"):
w.say(f" docmost B: cannot log in (http={code})"); return False
rc, code, out = w.app_curl(sub, "/api/spaces", "-b", jar, "-H", "Content-Type: application/json",
data="{}", method="POST")
ok = code in ("200", "201") and B["space"] in out
neg = "drillBnever" in out
w.say(f" docmost B: /api/spaces http={code} seeded-space-listed={ok} (negative control listed={neg})")
return ok and not neg
def break_edge(app, frm, to, port_from, port_to):
"""The failing edge: a REAL migrating image move, plus — in the DRILL template only — the
health probe pointed at a port the app does not answer. Both in one drill commit."""
fy = f"{w.DRILL}/templates/{app}/.felhom.yml"
f = open(fy).read()
m = re.search(r"(healthcheck:\n(?:.*\n){0,8}?\s+port: )" + str(port_from) + r"\b", f)
assert m, "probe port not found"
f = f[:m.end() - len(str(port_from))] + str(port_to) + f[m.end():]
open(fy, "w").write(f)
h = w.drill_bump(app, frm, to)
return h
def pg_state(container, db, user):
return w.guest(f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1; "
f"docker exec {container} psql -U {user} -d {db} -Atc \"select name from kysely_migration order by name desc limit 3\" 2>&1; "
f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from kysely_migration\" 2>&1")
def romm_seed_b(sub, A):
"""A SECOND RomM user, created by the first (admin) one — written after the backup."""
jar, tok = FX["romm"]._csrf(w, sub)
user = "drillb" + os.urandom(3).hex()
rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}",
"-u", f"{A['user']}:{A['pw']}", "-H", "Content-Type: application/json",
data=json.dumps({"username": user, "email": f"{user}@gate.invalid",
"password": "Drill-" + os.urandom(8).hex(), "role": "viewer"}),
method="POST")
w.say(f" romm seed B: POST /api/users (as the admin) http={code} {body[:120]}")
return {"user": user} if code in ("200", "201") else None
def romm_verify_b(sub, A, B):
jar, tok = FX["romm"]._csrf(w, sub)
rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}",
"-u", f"{A['user']}:{A['pw']}")
ok = code == "200" and B["user"] in body
neg = "drillbnever" in body
w.say(f" romm B: GET /api/users http={code} seeded-user-listed={ok} (negative control listed={neg})")
return ok and not neg
def tree(paths):
cmd = "; ".join(f"echo \"{p}: files=$(find {p} -type f 2>/dev/null | wc -l) sum=$(find {p} -type f -exec sha256sum {{}} + 2>/dev/null | sort | sha256sum | cut -c1-16)\"" for p in paths)
return w.guest(cmd)
def my_state(container="romm-db", db="romm"):
# The root password is used INSIDE the container from its own env — it never leaves it.
q = lambda sql: f"docker exec {container} sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" {db}' 2>&1 | tr '\\n' ' '"
return w.guest(f"""echo -n "tables=$({q("select count(*) from information_schema.tables where table_schema=database()")}) "
echo -n "alembic=$({q("select version_num from alembic_version")}) "
echo -n "users=$({q("select count(*) from users")})"
""")
def vik_seed_b(sub, A):
"""A second project, created AFTER the backup — plus a task with a real ATTACHMENT (a file the
app writes into its files volume), uploaded through the app's own attachment API."""
tok, why = FX["vikunja"]._token(w, sub, A)
title = "drillB-" + os.urandom(4).hex()
rc, code, body = w.app_curl(sub, "/api/v1/projects", "-H", f"Authorization: Bearer {tok}",
"-H", "Content-Type: application/json", data=json.dumps({"title": title}), method="PUT")
if code not in ("200", "201"):
w.say(f" vikunja B: project refused {code} {body[:120]}"); return None
pid = json.loads(body)["id"]
rc, code, body = w.app_curl(sub, f"/api/v1/projects/{pid}/tasks", "-H", f"Authorization: Bearer {tok}",
"-H", "Content-Type: application/json", data=json.dumps({"title": "task-" + title}), method="PUT")
tid = json.loads(body)["id"] if code in ("200", "201") else None
content = "drill attachment " + os.urandom(8).hex()
fn = "/tmp/vik-att.txt"; open(fn, "w").write(content)
rc, code2, body2 = w.app_curl(sub, f"/api/v1/tasks/{tid}/attachments", "-H", f"Authorization: Bearer {tok}",
"-F", f"files=@{fn}", method="PUT")
w.say(f" vikunja seed B: project http=200 task={tid} attachment upload http={code2} {body2[:120]}")
return {"title": title, "pid": pid, "tid": tid, "att": content}
def vik_verify_b(sub, A, B):
tok, why = FX["vikunja"]._token(w, sub, A)
if not tok:
w.say(f" vikunja B: cannot log in {why}"); return {"project": False, "attachment": False}
rc, code, body = w.app_curl(sub, f"/api/v1/projects/{B['pid']}", "-H", f"Authorization: Bearer {tok}")
proj = code == "200" and B["title"] in body
rc, code, body = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments", "-H", f"Authorization: Bearer {tok}")
att_ok = False
try:
atts = json.loads(body)
if atts:
aid = atts[0]["id"]
rc, c3, b3 = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments/{aid}", "-H", f"Authorization: Bearer {tok}")
att_ok = c3 == "200" and B["att"] in b3
except Exception as e:
w.say(f" vikunja B: attachments list unreadable http={code} {body[:120]}")
w.say(f" vikunja B: project readback={proj} attachment content readback={att_ok}")
return {"project": proj, "attachment": att_ok}
@@ -0,0 +1,14 @@
10:16:50 === prep vikunja: deploy at the drill FROM pin
10:16:50 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
10:16:55 [1] deployed, controller state=running, pinned={'vikunja': 'vikunja/vikunja:2.3.0'}
10:16:55 deployed: True
10:16:56 vikunja: register http=200
10:16:56 vikunja: create project http=201
10:16:56 vikunja: readback of the seeded project http=200 ok=True
10:16:56 C1 A: True
10:16:56 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'}
10:17:56 [4] backup idle; last=None
10:17:57 vikunja seed B: project http=200 task=1 attachment upload http=200 {"errors":null,"success":[{"id":1,"task_id":1,"created_by":{"id":1,"name":"","username":"drill6d84b2","created":"2026-09
10:17:57 vikunja B: project readback=True attachment content readback=True
10:17:57 B reads back: True
10:18:00 named volumes: ['vikunja_vikunja_data', 'vikunja_vikunja_db']
@@ -0,0 +1,9 @@
10:18:03 db state before the update: tables=36 ledger=117 newest=SCHEMA_INIT
10:18:05 image: vikunja/vikunja:2.3.0
10:18:12 VOL vikunja_vikunja_data bytes=4129 files=2 | copy bytes=4129 files=2 | cp -a 0.42s | tar 0.41s (2560 B)
VOL vikunja_vikunja_db bytes=2471792 files=4 | copy bytes=2471792 files=4 | cp -a 0.42s | tar 0.40s (2470912 B)
stop 0.18s copy(all, cp -a + the tar measurement) 5.03s start 0.23s
free on the docker root: 18.18 GiB
10:18:14 front door answering again 2.1s after the start
@@ -0,0 +1,2 @@
10:18:15 [5] drill commit 9772e7c0b9f5: vikunja vikunja/vikunja:2.3.0 -> vikunja/vikunja:2.6.0 (push rc=0)
10:18:20 badge caught up after 4.5 s
@@ -0,0 +1,7 @@
10:18:20 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
10:18:20 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None
10:18:21 + 1.0s phase=pulling label=Új verzió letöltése… err=None hold=None
10:18:24 + 4.1s phase=starting label=Indítás az új verzióval… err=None hold=None
10:18:25 + 5.1s phase=verifying label=Működés ellenőrzése… err=None hold=None
10:19:56 + 95.4s phase=failed label=A frissítés nem sikerült err=A(z) vikunja frissítése 2026-09-23 10:19-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:17 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza. hold=A(z) vikunja frissítése 2026-09-23 10:19-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:17 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.
10:19:56 {"final_phase": "failed", "hold_reason": "A(z) vikunja frissítése 2026-09-23 10:19-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:17 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.", "state": "stopped", "duration_s": 95.4}
@@ -0,0 +1,13 @@
10:19:56 === undo by F: vikunja
10:20:14 [INFO] [settings] restore hold CLEARED for vikunja
10:20:16 pinned_images:
vikunja: vikunja/vikunja:2.3.0
10:20:16 hold lift + pin back took 20.2 s (not part of a product undo)
10:20:19 volumes put back in 0.90s
10:20:19 product start -> 200
10:20:22 vikunja: readback of the seeded project http=200 ok=True
10:20:22 vikunja B: project readback=True attachment content readback=True
10:20:24 F RESULT: healthy=True after 2.7s A=True B=True db now: tables=36 ledger=117 newest=SCHEMA_INIT (before the update: tables=36 ledger=117 newest=SCHEMA_INIT)
@@ -0,0 +1,6 @@
10:20:25 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
10:20:25 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None
10:20:26 + 1.1s phase=pulling label=Új verzió letöltése… err=None hold=None
10:20:27 + 2.1s phase=verifying label=Működés ellenőrzése… err=None hold=None
10:21:58 + 93.3s phase=failed label=A frissítés nem sikerült err=A(z) vikunja frissítése 2026-09-23 10:21-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:19 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza. hold=A(z) vikunja frissítése 2026-09-23 10:21-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:19 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.
10:21:58 {"final_phase": "failed", "hold_reason": "A(z) vikunja frissítése 2026-09-23 10:21-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:19 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.", "state": "stopped", "duration_s": 93.3}
@@ -0,0 +1,8 @@
10:21:58 === undo by D: vikunja
10:22:16 [INFO] [settings] restore hold CLEARED for vikunja
10:22:18 pinned_images:
vikunja: vikunja/vikunja:2.3.0
10:22:18 hold lift + pin back took 20.1 s (not part of a product undo)
10:22:18 vikunja has no database server: the product wrote NO safety dump (R-641). D's copy for it is a volume tar at safety-dump time, which is F with extra steps — recorded, not re-run.
@@ -0,0 +1,11 @@
10:22:23 === cut-off copy: vikunja (largest volume vikunja_vikunja_db)
10:22:27 copy container exit=0 (137 = killed)
source bytes=2920872 cut copy bytes=2920872
finished-marker in the cut copy: 1 -> F refuses to swap
D: no safety dump exists for this app (R-641)
10:22:40 RETRY — 2.9 MB copied inside 0.02 s, so the first cut missed; kill with no delay:
10:22:48 try 1: container exit=0 source bytes=2920872 cut copy bytes=2920872 finished-marker=1
try 2: container exit=0 source bytes=2920872 cut copy bytes=2920872 finished-marker=1
try 3: container exit=0 source bytes=2920872 cut copy bytes=2920872 finished-marker=1
@@ -0,0 +1,493 @@
#!/usr/bin/env python3
"""walk.py — ONE app's full update walk on guest 9202, through the product's own endpoints.
EVIDENCE, NOT PRODUCT. It presses exactly the buttons a person presses:
POST /api/stacks/<n>/deploy · POST /api/backup/run · POST /api/sync · POST /api/stacks/rescan
POST /api/stacks/<n>/update · POST /api/stacks/<n>/remove
and reads GET /api/stacks/<n>. No controller code exists for it.
The walk, per `09` §6.4 and the update-night brief §4:
1 deploy from the DRILL catalog at the LIVE pin
2 seed through the app's OWN front door (R-156: never a volume, never SQL)
3 read the seed back <- control C1; a fixture that cannot prove itself proves nothing
4 „Mentés most"
5 commit the real one-step bump to the DRILL repo, sync, rescan, read the badge in BOTH languages
6 press the guarded Update, record every phase with timestamps
7 read the seed back through the front door
8 the four version observables side by side
9 write the verdict record in `09`'s JSON shape
`inconclusive` is a first-class verdict and is NEVER collapsed into `failed`.
"""
import argparse, json, os, re, subprocess, sys, time
from datetime import datetime, timezone
SC = "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/6e5a1a3b-6d8c-4ee1-bc3f-c555eb3f7578/scratchpad"
EV = "/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/undo-bakeoff-2026-09-23"
DRILL = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill"
BASE = "https://192.168.0.114"
HOSTHDR = "Host: felhom.enkisfelhom.hu"
DOMAIN = "enkisfelhom.hu"
HP = "demo-hp"
LOG = []
def say(*a):
line = " ".join(str(x) for x in a)
ts = datetime.now().strftime("%H:%M:%S")
print(f"{ts} {line}", flush=True)
LOG.append(f"{ts} {line}")
def sh(args, timeout=300, inp=None):
try:
return subprocess.run(args, capture_output=True, text=True, timeout=timeout, input=inp)
except (subprocess.TimeoutExpired, OSError) as e:
return subprocess.CompletedProcess(args, 124, "", f"{e}")
def guest(script, timeout=600):
"""Run a bash script inside guest 9202. Piped as a file — never as an argument (quoting)."""
r = sh(["ssh", "-o", "ConnectTimeout=20", "-o", "StrictHostKeyChecking=accept-new", HP,
"cat > /tmp/w.sh; pct push 9202 /tmp/w.sh /tmp/w.sh >/dev/null 2>&1; "
"pct exec 9202 -- bash /tmp/w.sh; rm -f /tmp/w.sh"],
timeout=timeout, inp=script)
return r.stdout or ""
def login():
pw = open(f"{SC}/.ctlpw").read().strip()
sh(["curl", "-sk", "-D", f"{SC}/hdr.txt", "-o", "/dev/null", "-H", HOSTHDR,
"-X", "POST", "--data-urlencode", f"password={pw}", f"{BASE}/login"])
h = open(f"{SC}/hdr.txt").read()
m = re.search(r"felhom_session=[A-Za-z0-9._-]+", h, re.I)
if not m:
sys.exit("login failed: no session cookie")
open(f"{SC}/sess.txt", "w").write(m.group(0))
r = sh(["curl", "-sk", "-L", "-H", HOSTHDR, "-H", f"Cookie: {m.group(0)}", f"{BASE}/"])
c = re.search(r'<meta name="csrf-token" content="([^"]+)"', r.stdout or "")
if not c:
sys.exit("login failed: no csrf token")
open(f"{SC}/csrf.txt", "w").write(c.group(1))
def ctl(method, path, data=None, raw=False, tries=2):
"""One controller API call. Re-logs in once on a 302/401 — the controller's session store is
in memory, so any controller restart during the night invalidates it silently."""
for attempt in range(tries):
sess = open(f"{SC}/sess.txt").read().strip()
csrf = open(f"{SC}/csrf.txt").read().strip()
args = ["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", "-w", "\n%{http_code}"]
if method != "GET":
args += ["-H", f"X-CSRF-Token: {csrf}", "-H", "Content-Type: application/json",
"-X", method]
if data is not None:
args += ["--data", json.dumps(data)]
args.append(f"{BASE}{path}")
r = sh(args)
body, _, code = (r.stdout or "").rpartition("\n")
if code.strip() in ("302", "401") and attempt + 1 < tries:
login()
continue
if raw:
return code.strip(), body
try:
return code.strip(), json.loads(body)
except Exception:
return code.strip(), {"_raw": body[:600]}
return code.strip(), {"_raw": body[:600]}
def page(path):
sess = open(f"{SC}/sess.txt").read().strip()
r = sh(["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", f"{BASE}{path}"])
return r.stdout or ""
def app_curl(sub, path, *extra, method=None, data=None, timeout=45):
"""A call to the APP's own front door on 9202 — the household's route, not ours."""
args = ["curl", "-sSk", "--max-time", str(timeout), "-H", f"Host: {sub}.{DOMAIN}",
"-w", "\n%{http_code}"]
if method:
args += ["-X", method]
if data is not None:
args += ["--data-binary", "@-"]
args += list(extra) + [f"{BASE}{path}"]
r = sh(args, timeout=timeout + 30, inp=data)
body, _, code = (r.stdout or "").rpartition("\n")
return r.returncode, code.strip(), body
def stack(name):
_, d = ctl("GET", f"/api/stacks/{name}")
return (d.get("data") or {}) if isinstance(d, dict) else {}
def wait_app(sub, path="/", want=("200", "302", "303", "401", "403"), tries=60, delay=5):
"""Settling says the container runs; this says the APP answers. Not the same thing."""
last = None
for _ in range(tries):
rc, code, _ = app_curl(sub, path, timeout=15)
last = (rc, code)
if rc == 0 and code in want:
return True
time.sleep(delay)
say(f" app never answered on {sub}{path} (last rc={last[0]} code={last[1]})")
return False
# ------------------------------------------------------------------ the walk
DRIVE = "/mnt/felhom-drives/scratch_hdd/userdata"
# What THIS run generated for a deploy, per app. Deploy secrets are ENCRYPTED AT REST in
# `app.yaml` (`ENC:…`), which is right and which means a fixture cannot read an app's admin
# password back off the box — the household sees it once. So the value the harness itself
# generated is kept here for the life of the run, and nowhere else.
GENERATED = {}
def deploy_values(name, sub):
"""Fill EVERY required deploy field the way the wizard would, by asking the box what this app
asks for — `GET /api/stacks/<n>/deploy-fields` — instead of assuming DOMAIN+SUBDOMAIN.
Measured 2026-09-21: three apps in one batch refused at the deploy with a correct 400 because
a required field was absent — `HDD_PATH` (navidrome, audiobookshelf) and an admin password
(grafana). The refusals happen BEFORE anything is created (`deploy.go:324`), which is the only
reason this was safe to discover by running it (live-probes rule).
A `path` field must name a directory that ALREADY EXISTS (`deploy.go:330`), so one is made on
the scratch drive first — the same act the drive browser performs for a household.
"""
code, d = ctl("GET", f"/api/stacks/{name}/deploy-fields")
fields = (((d.get("data") or {}).get("metadata") or {}).get("deploy_fields")) or []
values = {"DOMAIN": DOMAIN, "SUBDOMAIN": sub}
made = []
for f in fields:
ev, ty = f.get("env_var"), f.get("type")
if ev in values:
continue
# `type: password` is MANDATORY whatever `required` says — `deploy.go:305-312` refuses
# when the caller sends none, deliberately ("the user needs to know their password"),
# while `.felhom.yml` declares `required: false` and the API serves that verbatim. A
# caller that trusts the contract gets a 400. Measured tonight on grafana; filed.
if not f.get("required") and ty != "password":
continue # the controller generates the optional secrets itself
if ty == "path":
p = f"{DRIVE}/{name}"
values[ev] = p
made.append(p)
elif ty in ("secret", "password"):
import secrets as _s
values[ev] = "Drill-" + _s.token_hex(12)
GENERATED.setdefault(name, {})[ev] = values[ev]
elif f.get("default"):
values[ev] = f["default"]
else:
values[ev] = f"drill-{name}"
if made:
guest("mkdir -p " + " ".join(made) + "; ls -ld " + " ".join(made))
say(f" [1] made the drive paths this app requires: {made}")
extra = [k for k in values if k not in ("DOMAIN", "SUBDOMAIN")]
if extra:
say(f" [1] required fields filled beyond DOMAIN/SUBDOMAIN: {extra}")
return values
def deploy(name, sub, extra_values=None):
st = stack(name)
if st.get("deployed"):
say(f" [1] {name} already deployed — reusing")
return True
values = deploy_values(name, sub)
if extra_values:
values.update(extra_values)
code, d = ctl("POST", f"/api/stacks/{name}/deploy", {"values": values})
say(f" [1] deploy -> {code} {str(d)[:120]}")
if code != "202":
return False
# WAIT FOR `deployed`, NOT FOR `running`. Measured 2026-09-21 on tandoor: docker reported the
# container `healthy` while the controller's own state read `unhealthy` — a gate on `running`
# alone therefore times out on an app that is up. The state is RECORDED rather than required;
# the real gate is the fixture's own `wait_app`, which asks whether the APP answers.
seen = None
for _ in range(90):
time.sleep(5)
st = stack(name)
seen = st.get("state")
# `deployed` alone is NOT enough and `state` alone is NOT right. Measured 2026-09-21:
# tandoor reads `unhealthy` while serving (R-618), so gating on "running" hangs; and romm
# read `deployed=True, state=degraded, pinned_images=None` twenty seconds in, i.e. the
# deploy had not finished writing app.yaml. The PIN is the deploy's own completion mark
# (`runComposeDeploy` writes it), so that is what to wait for.
pins = (st.get("app_config") or {}).get("pinned_images")
if st.get("deployed") and pins and seen in ("running", "unhealthy", "degraded"):
say(f" [1] deployed, controller state={seen}, "
f"pinned={(st.get('app_config') or {}).get('pinned_images')}")
if seen != "running":
say(f" [1] NOTE: the controller's own state is {seen!r}, not 'running' — recorded, "
f"not treated as a failure; the fixture's front-door wait is the real gate")
return True
say(f" [1] never became deployed (last controller state={seen!r})")
return False
def backup_now(name):
code, d = ctl("POST", "/api/backup/run")
say(f" [4] „Mentés most\" -> {code} {str(d)[:160]}")
for _ in range(90):
time.sleep(5)
c2, s = ctl("GET", "/api/backup/status")
dd = s.get("data") or {}
if not dd.get("running", False):
say(f" [4] backup idle; last={dd.get('last_run') or dd.get('last_db_dump')}")
return True
say(" [4] backup still running after 7.5 min — carrying on")
return False
def drill_bump(app, frm, to, service_hint=None):
"""Commit the edge to the DRILL repo. catalog_since set by hand (the drill repo has no gates).
`frm`/`to` may be comma-separated lists of the SAME length: an app whose own version lives in
two images (adventurelog's backend and frontend) moves both in one edge, while its engine
sidecar stays where it is — `09` §3b Q3's rule is per SERVICE, and an app-half edge must move
every service that carries the app's own version and no others.
"""
comp = f"{DRILL}/templates/{app}/docker-compose.yml"
fy = f"{DRILL}/templates/{app}/.felhom.yml"
s = open(comp).read()
froms = [x.strip() for x in frm.split(",") if x.strip()]
tos = [x.strip() for x in to.split(",") if x.strip()]
if len(froms) != len(tos):
say(f" [5] from/to lists differ in length: {froms} vs {tos}")
return None
for f1, t1 in zip(froms, tos):
if f"image: {f1}" not in s:
say(f" [5] FROM ref not found in compose: {f1}")
return None
s = s.replace(f"image: {f1}", f"image: {t1}")
open(comp, "w").write(s)
f = open(fy).read()
today = datetime.now().strftime("%Y-%m-%d")
f = re.sub(r'^catalog_since:.*$', f'catalog_since: "{today}"', f, count=1, flags=re.M)
open(fy, "w").write(f)
sh(["git", "-C", DRILL, "add", "-A"])
sh(["git", "-C", DRILL, "commit", "-q", "-m", f"DRILL {app}: {frm} -> {to}"])
r = sh(["git", "-C", DRILL, "push", "-q", "origin", "main"], timeout=120)
h = sh(["git", "-C", DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip()
say(f" [5] drill commit {h}: {app} {frm} -> {to} (push rc={r.returncode})")
return h
def sync_rescan(expect_app=None, expect_ref=None, tries=12, delay=5):
"""Sync, rescan, and — when told what to expect — WAIT FOR THE BADGE TO CATCH UP.
R-607: `POST /api/sync` answers "nincs valtozas" while the catalog HAS moved, and
`catalog_images` stays stale until a separate rescan. Tonight showed the rescan alone is not
enough either: mealie's badge read "Naprakesz" seconds after its bump was pushed, and the
Update that followed moved nothing and still reported "Frissitve". So when the caller knows
which reference should appear, this polls for it and SAYS HOW LONG IT TOOK — which is the
NUMBER R-607 asks for and has never had.
"""
t0 = time.time()
ctl("POST", "/api/sync")
time.sleep(2)
ctl("POST", "/api/stacks/rescan")
time.sleep(2)
if not expect_app or not expect_ref:
return None
for i in range(tries):
cat = stack(expect_app).get("catalog_images") or {}
if expect_ref in cat.values():
waited = round(time.time() - t0, 1)
if i:
say(f" [sync] the badge needed {waited}s and {i+1} sync+rescan rounds to catch up "
f"to {expect_ref} — R-607's window, measured")
return waited
time.sleep(delay)
ctl("POST", "/api/sync")
time.sleep(1)
ctl("POST", "/api/stacks/rescan")
say(f" [sync] the badge NEVER caught up to {expect_ref} in {round(time.time()-t0,1)}s — "
f"catalog_images = {stack(expect_app).get('catalog_images')}")
return None
def badges(name):
out = {}
for lang, suffix in (("hu", ""), ("en", "?lang=en")):
h = page(f"/apps/{name}{suffix}")
m = re.findall(r'<span class="tag tag-[^"]*"[^>]*title="([^"]*)"[^>]*>([^<]*)<', h)
out[lang] = [{"title": a.strip(), "text": b.strip()} for a, b in m][:3]
return out
def press_update(name, poll=1.0, cap_s=1800):
code, d = ctl("POST", f"/api/stacks/{name}/update")
say(f" [6] Update -> {code} {str(d)[:220]}")
if code not in ("202", "200"):
return {"accepted": False, "http": code, "refusal": d, "phases": [], "duration_s": 0}
phases, seen, t0 = [], None, time.time()
while time.time() - t0 < cap_s:
st = stack(name)
ph = st.get("update_phase")
if ph != seen:
seen = ph
rec = {"t": round(time.time() - t0, 1), "phase": ph,
"label": st.get("update_phase_label"), "updating": st.get("updating"),
"error": st.get("update_error"), "hold": st.get("hold_reason")}
phases.append(rec)
say(f" +{rec['t']:>6.1f}s phase={ph} label={rec['label']} "
f"err={rec['error']} hold={rec['hold']}")
if not st.get("updating") and ph in ("done", "failed", None) and time.time() - t0 > 3:
break
time.sleep(poll)
st = stack(name)
return {"accepted": True, "http": code, "phases": phases,
"duration_s": round(time.time() - t0, 1),
"final_phase": st.get("update_phase"), "update_error": st.get("update_error"),
"hold_reason": st.get("hold_reason"), "state": st.get("state")}
def observables(name):
st = stack(name)
ac = st.get("app_config") or {}
live = guest(f"""
grep -E '^\\s+image:' /opt/docker/stacks/{name}/docker-compose.yml 2>/dev/null | sed 's/^ *//'
echo '---inspect---'
for c in $(docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'); do
echo -n "$c "; docker inspect "$c" --format '{{{{.Config.Image}}}} running={{{{.State.Running}}}} restarts={{{{.RestartCount}}}}'
done
""")
a, _, b = live.partition("---inspect---")
return {
"pinned_images": ac.get("pinned_images"),
"installed_images": {k: (v.get("ref") if isinstance(v, dict) else v)
for k, v in (ac.get("installed_images") or {}).items()},
"catalog_images": st.get("catalog_images"),
"live_compose_image_lines": [x for x in a.strip().splitlines() if x.strip()],
"docker_inspect": [x for x in b.strip().splitlines() if x.strip()],
}
def app_logs(name, lines=400):
"""The app's own container log, DECODED. The endpoint answers a JSON envelope whose `logs` is
one string with escaped newlines — a scan over the envelope sees a single enormous line and
finds nothing, which reads exactly like "the app printed no migration line" and is not. R-96
rule 3 in a new place: an absent line is not evidence when the instrument cannot see lines."""
code, d = ctl("GET", f"/api/stacks/{name}/logs?lines={lines}")
if isinstance(d, dict):
data = d.get("data")
if isinstance(data, dict) and isinstance(data.get("logs"), str):
return data["logs"]
if isinstance(d.get("_raw"), str):
return d["_raw"]
return str(d)
def write_verdict(rec, appdir):
os.makedirs(appdir, exist_ok=True)
p = os.path.join(appdir, "verdict.json")
json.dump(rec, open(p, "w"), indent=2, ensure_ascii=False)
say(f" [9] verdict {rec['verdict']} -> {p}")
def remove(name):
"""Remove through the PRODUCT, never `docker rm` (live-probes rule). The remove endpoint
refuses a running stack — `409 still running` — so the stop is part of the act, not a tidy-up."""
c1, d1 = ctl("POST", f"/api/stacks/{name}/stop")
say(f" [X] stop -> {c1} {str(d1)[:100]}")
for _ in range(24):
time.sleep(5)
if stack(name).get("state") != "running":
break
code, d = ctl("POST", f"/api/stacks/{name}/remove",
{"remove_hdd_data": True, "remove_backups": True})
say(f" [X] remove (with drive data) -> {code} {str(d)[:160]}")
if code == "409":
# R-442's fail-closed guard: when the storage subsystem cannot RESOLVE the app's drive
# path, the removal is REFUSED and the app is kept rather than half-deleted. On guest 9202
# `/api/disks` answers `agent not configured`, so every app deployed with an HDD_PATH hits
# this. The household's other choice — remove the app, KEEP the data — is accepted, and the
# harness takes it, then tidies its own directory by name at teardown.
say(" [X] refused because the drive path cannot be resolved (R-442, fail-closed and right)"
" — removing the app and KEEPING the drive data instead")
code, d = ctl("POST", f"/api/stacks/{name}/remove",
{"remove_hdd_data": False, "remove_backups": True})
say(f" [X] remove (keeping drive data) -> {code} {str(d)[:160]}")
time.sleep(5)
st = stack(name)
left = guest(f"ls -d /opt/docker/stacks/{name} 2>/dev/null; "
f"docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'")
say(f" [X] after remove: deployed={st.get('deployed')} leftovers={left.strip()!r}")
return code
def app_env(name, key):
"""Read one deploy value the CUSTOMER was given (e.g. the generated admin password) from the
app's own `app.yaml`. This is not seeding — it is how the household logs in; the controller
shows them the same value. Data still goes in through the app's own front door."""
out = guest(f"grep -E '^\\s*{key}:' /opt/docker/stacks/{name}/app.yaml 2>/dev/null | head -1")
if ":" in out:
return out.split(":", 1)[1].strip().strip('"').strip("'")
return ""
def snapshots(name):
"""The restorable copies the backups page offers for this app."""
code, d = ctl("GET", f"/api/backup/snapshots?stack={name}")
data = d.get("data") if isinstance(d, dict) else None
if isinstance(data, dict):
for k in ("snapshots", "items", "restore_points"):
if isinstance(data.get(k), list):
return data[k]
return data if isinstance(data, list) else []
def restore(name, snapshot_id=None, wait_s=1200):
"""The household's own way out: the „Visszaállítás a mentésből" button on the backups page.
A FORM post, not an API call — `POST /backup/restore` with `_csrf`, `stack_name`,
`snapshot_id` — because that is the button the sentence tells them to press.
"""
snaps = snapshots(name)
if snapshot_id is None:
if not snaps:
say(f" [R] no restorable copy offered for {name}")
return {"ok": False, "why": "no snapshot offered", "snapshots": snaps}
first = snaps[0]
snapshot_id = first.get("id") or first.get("snapshot_id") or first.get("short_id")
say(f" [R] restoring {name} from snapshot {snapshot_id!r} (of {len(snaps)} offered)")
sess = open(f"{SC}/sess.txt").read().strip()
csrf = open(f"{SC}/csrf.txt").read().strip()
r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", HOSTHDR, "-H", f"Cookie: {sess}",
"-X", "POST",
"--data-urlencode", f"_csrf={csrf}",
"--data-urlencode", f"stack_name={name}",
"--data-urlencode", f"snapshot_id={snapshot_id}",
f"{BASE}/backup/restore"], timeout=180)
head = (r.stdout or "").split("\n")[0].strip()
loc = [l for l in (r.stdout or "").split("\n") if l.lower().startswith("location:")]
say(f" [R] POST /backup/restore -> {head} {loc[:1]}")
t0 = time.time()
last = None
while time.time() - t0 < wait_s:
code, d = ctl("GET", "/api/backup/restore-status")
dd = d.get("data") or {}
cur = (dd.get("running"), dd.get("phase") or dd.get("state"), dd.get("message"))
if cur != last:
say(f" +{round(time.time()-t0,1):>6.1f}s restore {cur}")
last = cur
if not dd.get("running", False) and time.time() - t0 > 5:
break
time.sleep(2)
st = stack(name)
say(f" [R] after restore: state={st.get('state')} hold={st.get('hold_reason')!r} "
f"phase={st.get('update_phase')}")
return {"ok": True, "snapshot_id": snapshot_id, "snapshots": snaps,
"http": head, "location": loc[:1], "seconds": round(time.time() - t0, 1),
"state_after": st.get("state"), "hold_after": st.get("hold_reason"),
"observables_after": observables(name)}
+1
View File
@@ -814,6 +814,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-642** | **[P3-LOW] `POST /api/stacks/{name}/start` answers 200 *Stack … start completed* while the app is crash-looping.** MEASURED 2026-09-23 on 9202 twice: docmost 0.95.0 and romm 5.0.0 started on data their newer versions had migrated — both refused and restarted in a loop (`Restarting (1)`, front door 404) behind a 200. The same false-green class as R-443 (closed for the Update) and R-635, on the Start. It matters now because an undo (R-637) must never read the start's return as success. Evidence: `audits/update-rulings-2026-09-23/README.md`, `docmost-43`, `romm-43`. | **OPEN — P3; owner: CC** |
| **R-643** | **[P2-MEDIUM] The ruled chain leaves the automatic update leg AT MOST 15 MINUTES a night.** FOUND 2026-09-23 while writing the build plan for `09` §3 decision 11 (*updates after the off-site copy, before the full-system backup*). The off-site leg starts at W+105m (`cmd/controller/main.go:961`) and the full-system backup's gate opens at W+2h (`quiesce/quiesce.go:656`, span to W+6h); the legs are clock-scheduled, not chained. One step takes ~1 min when it works and ~2–6 min when it fails and is undone. Options and the recommendation (the full-system backup waits for the leg inside its own window; the leg stops starting steps at W+5h) are in `09` §6.4. | **WAITING-ON-OPERATOR — `09` §6.4's one open point; owner: CC once answered** |
| **R-644** | **[P3-LOW] `gokapi` on scratch guest 9202 is crash-looping — 329 restarts by 2026-09-23 07:51 UTC, *password does not appear to be a SHA-1 hash* — and the controller still lists it deployed.** OBSERVED at the start of the 2026-09-23 session, not caused by it. The twenty-eight walk's teardown (2026-09-22) removed a `gokapi` container left by R-633 by name; a `gokapi` is running again, recorded `deployed: true`. Not investigated (scope). Likely the R-633/R-634 shape — a restore-then-remove race leaving a record — and a scratch-box fact, not a customer one; filed so the next drill does not read it as its own. | **OPEN — P3; owner: CC; investigate before the next drill on 9202** |
| **R-645** | **[P3-LOW] Lifting an update hold by the operator CLI lets the recovery unit be re-captured with the FAILED new definition within seconds — the copy the hold sentence names is overwritten.** MEASURED 2026-09-23 on 9202 during the undo bake-off: docmost was held at 07:59:16Z after a failed 0.95.0 → 0.96.0 update; `--clear-restore-hold docmost` + the controller restart it requires ran at ~07:59:23Z, and at **07:59:26Z** the controller logged *Recovery unit captured for docmost* — the unit's `compose/docker-compose.yml` now named `docmost/docmost:0.96.0`, the version that had just failed. The hold sentence had pointed the household at that unit („saját meghajtó, … 09:55"). The hold is what keeps the nightly legs off a held app (`isHeld`, v0.238.1); once it is lifted by hand, the checksum-gated refresh sees a changed definition and captures it. **Who it hits:** an operator who lifts a hold to inspect or repair, before restoring. With the undo (R-637) a failed update no longer holds unless the undo also fails, so the path is rarer — it does not go away. Candidate shapes, none chosen: the CLI refuses to lift an UPDATE hold (only a restore lifts it); or the lift also puts the pin back; or the capture skips an app whose pin is not what it is running. Evidence: `audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt` (the invalid run) and README §"Three things". | **OPEN — P3; owner: CC** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.
Clearing a row means the check was DONE and its result recorded in that R-row —