Undo bake-off: copy the folder wins (09 §3 decisions 19-20, §6.1a)
gates / gates (push) Successful in 26s
gates / gates (push) Successful in 26s
- 09 §3: decision 19 (the copy method is chosen by a bake-off) and 20 (the full-system backup waits for the update leg, inside its window; built later). - Bake-off on 9202, docmost / romm / vikunja: both methods pass every case; the folder copy wins because an app with no database server gets no dump, so dump-and-load would need the folder copy anyway. 1-5 s extra downtime, ~420 MB/s, disk = the volumes. - R-645 filed: lifting an update hold by hand lets the recovery unit be re-captured with the failed definition within seconds. Documents and evidence only; product code follows in the controller. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -265,6 +265,21 @@ builds them (`audits/update-rulings-2026-09-23/`).
|
||||
18. **Fleet view** (Q7). The report carries, per compose service (so the database too), the installed
|
||||
reference, the catalog reference and the badge state. **Built later**, when the fleet grows.
|
||||
|
||||
### 2026-09-23 (afternoon) — two more operator rulings
|
||||
|
||||
19. **The undo's copy method is chosen by a bake-off, not by assumption.** Two methods are measured on
|
||||
the same apps — *dump and load* (the database as a text file, loaded back) and *copy the folder*
|
||||
(the app's named volumes copied with its containers stopped, put back on failure). The simpler one
|
||||
that passes every case is built; if both pass, the folder copy wins on simplicity unless its
|
||||
downtime or disk cost fails the bake-off's limits. **Why:** the morning spike found two traps in the
|
||||
dump route (a cut-off file loads as success; a failed MariaDB load is half old, half new) and the
|
||||
folder route had not been measured at all. Result: §6.1a.
|
||||
|
||||
20. **On update nights, the full-system backup waits for the update leg, inside its own window**
|
||||
(answers R-643 / §6.4 part 7's open point). The leg stops starting new steps at **W+5h**, so the
|
||||
full-system backup keeps at least one hour of its [W+2h, W+6h) window. **Built with §6.4 part 7,
|
||||
not before.**
|
||||
|
||||
**RomM follow-ups, operator-agreed the same day:** the test bench watches memory after an update
|
||||
(`upgrade-test.py`, 2026-09-23); a version move checks the memory limit (gate or checklist — §6.4);
|
||||
R-636's louder repeated alarm.
|
||||
@@ -798,6 +813,19 @@ product path that loads a safety dump back exists (`rollbackSafetyDump`), but on
|
||||
restore calls it. The eight things the build must add are listed in the audit; §6.4 part 1 prices
|
||||
them.
|
||||
|
||||
**THE COPY METHOD, chosen by the bake-off of 2026-09-23 afternoon (decision 19):
|
||||
COPY THE FOLDER.** `audits/undo-bakeoff-2026-09-23/README.md`. Both methods passed every case on all
|
||||
three apps (seeds before and after the backup, ledger equal, a cut-off copy caught before anything is
|
||||
swapped or loaded). The folder copy wins because an app with no database server gets no dump at all,
|
||||
so the dump route would have needed the folder copy anyway. **Shape:** after the pull and just before
|
||||
`up` — where the app is stopped anyway to be recreated — each NAMED volume the app owns is copied with
|
||||
`cp -a` into a sibling volume `<volume>.pre-update-<stamp>` by a helper container, which writes a
|
||||
finished-marker last; a bind-mounted user folder is never copied and never touched. On a failed health
|
||||
check the copy is validated (marker present) and put back, the old definition and pin are restored
|
||||
from the journal's own copies, and the old version is checked with the OLD `.felhom.yml` probe.
|
||||
Measured cost: ≈ 1–5 s extra downtime for the three apps, ≈ 420 MB/s, disk = the volumes' size (the
|
||||
update refuses before moving anything when the copy would breach the 2 GB floor).
|
||||
|
||||
**The release could not reach the fleet by floor — R-472.** The hub holds a controller floor above the
|
||||
vouched golden (publish-train rule 1), so under the weekly golden cadence (R-468) v0.237.0 and v0.238.0
|
||||
were hand-deployed to the demo guests. **RESOLVED by §3 decision 7 (hub v0.112.0):** v0.239.0 reached
|
||||
|
||||
@@ -0,0 +1,10 @@
|
||||
git:
|
||||
branch: main
|
||||
repo_url: https://gitea.dooplex.hu/admin/app-catalog-drill.git
|
||||
sync_interval: 15m
|
||||
token: <redacted>
|
||||
username: "admin"
|
||||
hub:
|
||||
update:
|
||||
health_timeout: 90s
|
||||
|
||||
@@ -0,0 +1,10 @@
|
||||
sync 200 {'ok': True, 'data': {'ok': True, 'message': 'Sablonok naprakészek — nincs változás'}, 'message': 'Sablonok na
|
||||
control 1 (9202 follows drill): a837c3a DRILL: FROM states for the undo bake-off
|
||||
image: docmost/docmost:0.95.0
|
||||
image: vikunja/vikunja:2.3.0
|
||||
image: rommapp/romm:5.0.0
|
||||
|
||||
control 2 (9201 on live): cfcfe52 upgrade-test: watch memory after the readback (harness v2, R-635/R-462)
|
||||
|
||||
control 3 (live main): cfcfe527842865aa35a6d2ae9d361872e36afc9a refs/heads/main
|
||||
|
||||
@@ -0,0 +1,8 @@
|
||||
10:22:59 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/nextcloud']
|
||||
10:22:59 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['NEXTCLOUD_ADMIN_PASSWORD', 'HDD_PATH']
|
||||
10:22:59 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
|
||||
10:23:44 [1] deployed, controller state=running, pinned={'nextcloud': 'nextcloud:34.0.1-apache', 'nextcloud-db': 'mariadb:12.3', 'nextcloud-redis': 'redis:7-alpine'}
|
||||
10:23:44 deployed: True
|
||||
10:23:48 nextcloud: occ user:add :: The account "drilladbeb4" was created successfully Display name set to "drilladbeb4"
|
||||
10:23:51 nextcloud: seeded user drilladbeb4
|
||||
10:23:51 seeded: True
|
||||
@@ -0,0 +1,5 @@
|
||||
10:26:02 300 WebDAV uploads through the front door: {'201': 300}
|
||||
10:26:03 PROPFIND lists 300 drill files (readback through the app); http 207
|
||||
10:26:06 449
|
||||
185374751 /v
|
||||
|
||||
@@ -0,0 +1,13 @@
|
||||
10:26:49 == D's cost on the same database: mariadb-dump while running
|
||||
dump 0.44s 1062572 B
|
||||
== F on the same database: stop, copy, start
|
||||
stop 1.58s
|
||||
cp -a 185 MB: 0.81s
|
||||
tar 185 MB: 0.53s
|
||||
start 0.63s
|
||||
== a 2 GiB synthetic volume (random bytes, not app data) for the extrapolation
|
||||
page cache could NOT be dropped (LXC) - warm-ish numbers
|
||||
cp -a 2 GiB: 4.87s
|
||||
tar 2 GiB: 4.55s
|
||||
/dev/loop1 69G 48G 18G 74% /var/lib/docker
|
||||
|
||||
@@ -0,0 +1,16 @@
|
||||
10:27:13 [X] stop -> 200 {'ok': True, 'message': 'Stack nextcloud stop completed'}
|
||||
10:27:18 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/nextcloud tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó
|
||||
10:27:18 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead
|
||||
10:27:46 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'nextcloud', 'volumes_removed': ['nextcloud_nextcloud_db_data', 'nextcloud_nextcloud_html', 'nextcloud_nextcloud_redis_data'],
|
||||
10:27:53 [X] after remove: deployed=False leftovers='/opt/docker/stacks/nextcloud'
|
||||
10:27:56 removed docmost_docmost_postgres_data.pre-undo
|
||||
removed docmost_docmost_redis_data.pre-undo
|
||||
removed docmost_docmost_storage.pre-undo
|
||||
removed romm_romm_config.pre-undo
|
||||
removed romm_romm_db_data.pre-undo
|
||||
removed romm_romm_redis_data.pre-undo
|
||||
removed vikunja_vikunja_data.pre-undo
|
||||
removed vikunja_vikunja_db.pre-undo
|
||||
nextcloud drive folder removed
|
||||
0
|
||||
|
||||
@@ -0,0 +1,59 @@
|
||||
# The undo bake-off, 2026-09-23 — copy the folder (F) against dump and load (D)
|
||||
|
||||
`09` §3 decision 19. Venue: scratch guest **9202**, controller v0.262.1, drill catalog (reset to live
|
||||
`cfcfe5278428` first; `a837c3a7d8bd` FROM states), `update.health_timeout: 90s`. Same three apps and
|
||||
edges as the morning spike (`../update-rulings-2026-09-23/`), each a real migrating edge made to fail
|
||||
a deliberately wrong probe in the DRILL template only. Seed A before the backup, seed B after it,
|
||||
both through the app's own front door. Driver: `bakeoff.py` (stages per app: prep → fcopy → break →
|
||||
update → undoF → update → undoD → cutoff). Time-boxed at 2 h; took ~40 min.
|
||||
|
||||
**Method, stated so it is not over-read.** Both undos were done BY HAND in the order the product
|
||||
would take them. The hold was lifted with the operator CLI + a controller restart (the only exit that
|
||||
exists) and the product's own Start supplied the app's secrets.
|
||||
|
||||
## The table
|
||||
|
||||
| criterion | limit | **F** docmost / romm / vikunja | **D** docmost / romm / vikunja |
|
||||
|---|---|---|---|
|
||||
| seed A and seed B read back after the undo | both, every app | **yes·yes / yes·yes / yes·yes** | yes·yes / yes·yes / — (no safety dump exists: R-641) |
|
||||
| tables / ledger equal to before the update | equal | **42 · 48 = / alembic 0095 = ¹ / 36 · 117 =** | 42 · 48 = / **24 base tables** · 0095 = / — |
|
||||
| a cut-off copy detected before anything is swapped or loaded | detected | **docmost 4 of 4 cuts** (container exit 137, 42.9–68.4 of 69.2 MB copied, no finished-marker); **romm** (105.8 of 160.2 MB, no marker); vikunja **could not be cut** — 2.9 MB completes before a kill lands ² | docmost and romm: the half file lacks `-- PostgreSQL database dump complete` / `-- Dump completed` → refused; vikunja — |
|
||||
| extra downtime before the update | ≤ 30 s | stop + copy: **≈ 2.4 s / ≈ 5.3 s / ≈ 1 s** (stop 0.52 / 3.79 / 0.18 s; copies 0.4–1.0 s per volume incl. the helper's start) | 0 — the dump runs against the running database |
|
||||
| rate | stated | cp -a **185 MB (Nextcloud's MariaDB) in 0.81 s; 2 GiB in 4.87 s ≈ 420 MB/s**; tar 0.53 s / 4.55 s. Page cache could NOT be dropped in the guest, so these are warm-ish. **A 5 GB database ≈ 12 s at this rate**; on a cold cache or a spinning system disk expect 25–50 s | `mariadb-dump` of the same Nextcloud: 0.44 s, 1.06 MB |
|
||||
| disk needed | stated; refuse near the 2 GB floor | the app's named volumes: **51 MB / 175 MB / 2.5 MB** (Nextcloud 185 MB) | the dump: 136 KB / 63 KB / — (Nextcloud 1 MB) |
|
||||
| product code the build needs | estimated | **≈ 300 lines**, one mechanism for every app class | ≈ 450: the marker, two engine-specific empty-then-load paths, **and F's own machinery anyway** for apps with no database server |
|
||||
|
||||
¹ romm's base-table count was not captured on the F run (a quoting bug in the evidence script, fixed
|
||||
before the D run, which read 24). F puts the whole datadir back byte-for-byte (copy bytes = source
|
||||
bytes, file counts equal, `romm-2-fcopy.txt`), and the migration ledger read back equal.
|
||||
² The marker logic is the same one proven on the other two; a copy that finishes cannot be cut.
|
||||
|
||||
## The choice — F, copy the folder (decision 19)
|
||||
|
||||
**Both methods pass every row.** Decision 19 then says the folder copy wins on simplicity unless its
|
||||
downtime or disk cost fails the limits — and neither does: ≤ 5.3 s for these three, and the disk
|
||||
cost is refused near the 2 GB floor by construction. **The decisive fact is the third column:** an
|
||||
app with no database server gets no safety dump, so D would have to carry F's volume copy anyway.
|
||||
F is one mechanism; D is F plus two engine loaders.
|
||||
|
||||
**Copy tool: `cp -a` into a sibling volume** (`<volume>.pre-update-<stamp>`), not `tar`: the two ran at
|
||||
the same speed, and a sibling volume puts the undo back with the same `cp -a` and needs no space on
|
||||
the app's own drive.
|
||||
|
||||
## Three things the bake-off found that the build must respect
|
||||
|
||||
1. **The undo must keep its own copy of the old definition.** Lifting the hold made the box re-capture
|
||||
docmost's recovery unit ten seconds later — with the NEW definition — and a pin-back that read the
|
||||
unit started the new version on the restored data, which migrated again (`docmost-40-undoF.txt`,
|
||||
invalid run, kept as evidence; redo `docmost-41`). R-639, seen live.
|
||||
2. **Killing `docker run` does not stop the copy** — the container runs on and finishes
|
||||
(`docmost-60-cutoff.txt`, first block). The product must judge a copy by the helper container's own
|
||||
exit status and the finished-marker it writes last, never by the client.
|
||||
3. **Where the copy is taken decides the downtime.** Taken after the pull and just before `up` — where
|
||||
the update stops the app anyway to recreate it — the extra downtime is the copy alone.
|
||||
|
||||
## Evidence
|
||||
|
||||
`docmost-*`, `romm-*`, `vikunja-*` (per stage), `70`–`73` (the rate test on Nextcloud, removed after),
|
||||
`01`/`02` (repoint + three controls). Nextcloud and every `.pre-undo`/`.cut`/`.rate` volume were
|
||||
removed by name at the end of the phase (`73-rate-teardown.txt`).
|
||||
@@ -0,0 +1,266 @@
|
||||
#!/usr/bin/env python3
|
||||
"""bakeoff.py — `09` §3 decision 19: the two ways of keeping the undo's last-second copy, measured on
|
||||
the same three apps (docmost / PostgreSQL, romm / MariaDB, vikunja / SQLite in a volume), guest 9202,
|
||||
drill catalog.
|
||||
|
||||
F — copy the folder: the app's NAMED volumes copied with its containers stopped; put back on failure.
|
||||
D — dump and load, as fixed by the morning spike: completion marker first; PostgreSQL drops and
|
||||
recreates the dump's schemas inside the load's transaction; MariaDB drops every table, then loads.
|
||||
|
||||
EVIDENCE, NOT PRODUCT. Product acts go through the endpoints the UI invokes (deploy, backup, sync,
|
||||
rescan, update, start, remove). The undo has no product path yet, so it is done by hand — the hold is
|
||||
lifted with the operator CLI + a controller restart (the only exit that exists), and the product's own
|
||||
Start supplies the app's secrets (they are never decrypted here).
|
||||
|
||||
Usage: python3 bakeoff.py <stage> <app> stages: prep fcopy break update undoF undoD cutoff state
|
||||
"""
|
||||
import json, os, sys, time
|
||||
sys.path.insert(0, ".")
|
||||
import walk as w
|
||||
from spike import FX, SUB, load, save, ts, break_edge, docmost_seed_b, docmost_verify_b, romm_seed_b, \
|
||||
romm_verify_b, vik_seed_b, vik_verify_b
|
||||
|
||||
EDGE = {"docmost": ("docmost/docmost:0.95.0", "docmost/docmost:0.96.0", 3000, 3999),
|
||||
"romm": ("rommapp/romm:5.0.0", "rommapp/romm:5.3.0", 8080, 8999),
|
||||
"vikunja": ("vikunja/vikunja:2.3.0", "vikunja/vikunja:2.6.0", 3456, 3999)}
|
||||
UNIT = {"docmost": "/mnt/sys_drive/felhom-data/backups/primary/docmost",
|
||||
"vikunja": "/mnt/sys_drive/felhom-data/backups/primary/vikunja",
|
||||
"romm": "/mnt/felhom-drives/scratch_hdd/userdata/romm/backups/primary/romm"}
|
||||
SUFFIX = ".pre-undo"
|
||||
|
||||
|
||||
def seed_b(app, sub, A):
|
||||
return {"docmost": docmost_seed_b, "romm": romm_seed_b, "vikunja": vik_seed_b}[app](sub, A)
|
||||
|
||||
|
||||
def verify_b(app, sub, A, B):
|
||||
r = {"docmost": docmost_verify_b, "romm": romm_verify_b, "vikunja": vik_verify_b}[app](sub, A, B)
|
||||
return all(r.values()) if isinstance(r, dict) else r
|
||||
|
||||
|
||||
def db_state(app):
|
||||
"""The table count and the migration ledger, asked of the engine (or the SQLite file, read-only)."""
|
||||
if app == "docmost":
|
||||
return w.guest("docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1 | tr '\\n' ' '; "
|
||||
"docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*)||' ledger, newest '||max(name) from kysely_migration\" 2>&1").strip()
|
||||
if app == "romm":
|
||||
q = lambda sql: f"docker exec romm-db sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" romm' 2>&1 | tr '\\n' ' '"
|
||||
base_tables = q("select count(*) from information_schema.tables where table_schema=database() and table_type=0x42415345205441424C45")
|
||||
alembic = q("select version_num from alembic_version")
|
||||
return w.guest(f'echo "tables=$({base_tables}) alembic=$({alembic})"').strip()
|
||||
return w.guest("""python3 - <<'PY'
|
||||
import sqlite3
|
||||
c=sqlite3.connect('file:/var/lib/docker/volumes/vikunja_vikunja_db/_data/vikunja.db?mode=ro',uri=True)
|
||||
t=c.execute("select count(*) from sqlite_master where type='table'").fetchone()[0]
|
||||
m=c.execute("select count(*), max(id) from migration").fetchone()
|
||||
print(f"tables={t} ledger={m[0]} newest={m[1]}")
|
||||
PY""").strip()
|
||||
|
||||
|
||||
def volumes(app):
|
||||
return [v for v in w.guest(f"docker volume ls -q --filter label=com.docker.compose.project={app}").split()
|
||||
if not v.endswith(SUFFIX)]
|
||||
|
||||
|
||||
def containers(app):
|
||||
return w.guest(f"docker ps -a --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'").split()
|
||||
|
||||
|
||||
def stage_prep(app):
|
||||
s = {"app": app, "sub": SUB[app]}
|
||||
w.say(f"=== prep {app}: deploy at the drill FROM pin")
|
||||
w.say("deployed:", w.deploy(app, s["sub"]))
|
||||
s["seedA"] = FX[app].seed(w, s["sub"], w.say)
|
||||
w.say("C1 A:", FX[app].verify(w, s["sub"], s["seedA"], w.say))
|
||||
w.backup_now(app)
|
||||
w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60)
|
||||
s["seedB"] = seed_b(app, s["sub"], s["seedA"])
|
||||
w.say("B reads back:", verify_b(app, s["sub"], s["seedA"], s["seedB"]))
|
||||
s["volumes"] = volumes(app)
|
||||
w.say("named volumes:", s["volumes"])
|
||||
save(app, s)
|
||||
|
||||
|
||||
def stage_fcopy(app):
|
||||
"""F at safety-dump time: stop, copy every named volume (cp -a into a sibling volume, and
|
||||
separately a tar, for the rate), start. Measures the EXTRA downtime and the disk."""
|
||||
s = load(app)
|
||||
s["state_before_update"] = db_state(app)
|
||||
w.say("db state before the update:", s["state_before_update"])
|
||||
vols = s["volumes"]
|
||||
w.say(w.guest(f"cp /opt/docker/stacks/{app}/docker-compose.yml /root/pre-{app}-compose.yml; grep -m1 'image: {EDGE[app][0]}' /root/pre-{app}-compose.yml"))
|
||||
script = f"""
|
||||
set -u
|
||||
cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}')
|
||||
t0=$(date +%s.%N)
|
||||
docker stop $cs >/dev/null
|
||||
t1=$(date +%s.%N)
|
||||
for v in {' '.join(vols)}; do
|
||||
docker volume rm -f "$v{SUFFIX}" >/dev/null 2>&1
|
||||
docker volume create "$v{SUFFIX}" >/dev/null
|
||||
a=$(date +%s.%N)
|
||||
docker run --rm -v "$v":/from:ro -v "$v{SUFFIX}":/to alpine:3.20 sh -c 'cp -a /from/. /to/ && sync' || echo "COPY FAILED $v"
|
||||
b=$(date +%s.%N)
|
||||
bytes=$(docker run --rm -v "$v":/from:ro alpine:3.20 du -sb /from | cut -f1)
|
||||
files=$(docker run --rm -v "$v":/from:ro alpine:3.20 sh -c 'find /from | wc -l')
|
||||
cbytes=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 du -sb /to | cut -f1)
|
||||
cfiles=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 sh -c 'find /to | wc -l')
|
||||
mkdir -p /root/bk; c=$(date +%s.%N)
|
||||
docker run --rm -v "$v":/vol:ro -v /root/bk:/out alpine:3.20 tar cf /out/$v.tar -C /vol . ; d=$(date +%s.%N)
|
||||
tb=$(stat -c %s /root/bk/$v.tar); rm -f /root/bk/$v.tar
|
||||
python3 -c "print('VOL $v bytes=$bytes files=$files | copy bytes=$cbytes files=$cfiles | cp -a %.2fs | tar %.2fs (%s B)' % ($b-$a, $d-$c, '$tb'))"
|
||||
done
|
||||
t2=$(date +%s.%N)
|
||||
docker start $cs >/dev/null
|
||||
t3=$(date +%s.%N)
|
||||
python3 -c "print('stop %.2fs copy(all, cp -a + the tar measurement) %.2fs start %.2fs' % ($t1-$t0, $t2-$t1, $t3-$t2))"
|
||||
df -B1 --output=avail /var/lib/docker | tail -1 | awk '{{printf "free on the docker root: %.2f GiB\\n", $1/2^30}}'
|
||||
"""
|
||||
out = w.guest(script, timeout=1800)
|
||||
w.say(out)
|
||||
t0 = time.time()
|
||||
w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2)
|
||||
w.say(f"front door answering again {round(time.time()-t0,1)}s after the start")
|
||||
s["fcopy"] = out
|
||||
s["fcopy_at"] = ts()
|
||||
save(app, s)
|
||||
|
||||
|
||||
def stage_break(app):
|
||||
s = load(app)
|
||||
frm, to, p0, p1 = EDGE[app]
|
||||
s["break_commit"] = break_edge(app, frm, to, p0, p1)
|
||||
w.say("badge caught up after", w.sync_rescan(app, to), "s")
|
||||
save(app, s)
|
||||
|
||||
|
||||
def stage_update(app):
|
||||
s = load(app)
|
||||
r = w.press_update(app, poll=1.0)
|
||||
w.say(json.dumps({k: r[k] for k in ("final_phase", "hold_reason", "state", "duration_s")}, ensure_ascii=False))
|
||||
s.setdefault("updates", []).append(r)
|
||||
save(app, s)
|
||||
|
||||
|
||||
LIFT = """docker exec felhom-controller /usr/local/bin/felhom-controller --clear-restore-hold {app} 2>&1 | grep -E 'CLEARED|no restore hold'
|
||||
docker restart felhom-controller >/dev/null; sleep 15"""
|
||||
|
||||
# The OLD definition comes from the bake-off's OWN pre-update copy (/root/pre-<app>-compose.yml, taken
|
||||
# in fcopy), NEVER from the recovery unit: measured 2026-09-23 10:59, the unit was re-captured 10 s
|
||||
# after the hold was lifted — with the NEW definition — and a pin-back that read it started the new
|
||||
# version again. That is R-639 seen live, and why the product undo keeps its own copies.
|
||||
PINBACK = """S=/opt/docker/stacks/{app}; P=/root/pre-{app}-compose.yml
|
||||
grep -m1 'image: .*{frm}' $P >/dev/null || {{ echo "PRE-UPDATE COPY MISSING OR WRONG: $P"; exit 1; }}
|
||||
cp $P $S/applied-compose.yml; cp $P $S/docker-compose.yml
|
||||
sed -i 's#^\\(\\s*{svc}: \\){to}$#\\1{frm}#' $S/app.yaml; sed -n '/^pinned_images:/,$p' $S/app.yaml | head -4"""
|
||||
|
||||
|
||||
def lift_and_pinback(app):
|
||||
frm, to, _, _ = EDGE[app]
|
||||
t = time.time()
|
||||
w.say(w.guest(LIFT.format(app=app)))
|
||||
w.say(w.guest(PINBACK.format(unit=UNIT[app], app=app, svc=app, frm=frm, to=to)))
|
||||
w.login()
|
||||
return round(time.time() - t, 1)
|
||||
|
||||
|
||||
def stage_undoF(app):
|
||||
s = load(app)
|
||||
w.say(f"=== undo by F: {app}")
|
||||
w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)")
|
||||
vols = s["volumes"]
|
||||
script = f"""
|
||||
t0=$(date +%s.%N)
|
||||
cs=$(docker ps -a --filter label=com.docker.compose.project={app} -q); [ -n "$cs" ] && docker stop $cs >/dev/null
|
||||
for v in {' '.join(vols)}; do
|
||||
docker run --rm -v "$v{SUFFIX}":/from:ro -v "$v":/to alpine:3.20 sh -c 'rm -rf /to/..?* /to/.[!.]* /to/* ; cp -a /from/. /to/ && sync' || echo "RESTORE FAILED $v"
|
||||
done
|
||||
python3 -c "import time;print('volumes put back in %.2fs' % (time.time()-$t0))"
|
||||
"""
|
||||
w.say(w.guest(script, timeout=1800))
|
||||
t0 = time.time()
|
||||
c, d = w.ctl("POST", f"/api/stacks/{app}/start")
|
||||
w.say("product start ->", c)
|
||||
up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2)
|
||||
th = round(time.time() - t0, 1)
|
||||
A = FX[app].verify(w, s["sub"], s["seedA"], w.say)
|
||||
B = verify_b(app, s["sub"], s["seedA"], s["seedB"])
|
||||
st = db_state(app)
|
||||
w.say(f"F RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})")
|
||||
s["undoF"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st}
|
||||
save(app, s)
|
||||
|
||||
|
||||
def stage_undoD(app):
|
||||
"""D, on the SECOND failed update: its safety dump (written by the product, phase 3) is the copy."""
|
||||
s = load(app)
|
||||
w.say(f"=== undo by D: {app}")
|
||||
w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)")
|
||||
if app == "vikunja":
|
||||
w.say("vikunja has no database server: the product wrote NO safety dump (R-641). D's copy for it "
|
||||
"is a volume tar at safety-dump time, which is F with extra steps — recorded, not re-run.")
|
||||
return
|
||||
c, d = w.ctl("POST", f"/api/stacks/{app}/start") # the product supplies the env; the app half will refuse
|
||||
time.sleep(8)
|
||||
if app == "docmost":
|
||||
load_cmd = """{ echo 'DROP SCHEMA public CASCADE; CREATE SCHEMA public;'; cat "$D"; } | docker exec -i docmost-postgres psql -v ON_ERROR_STOP=1 --single-transaction -U docmost -d docmost > /root/dload.out 2>&1"""
|
||||
marker = "-- PostgreSQL database dump complete"
|
||||
else:
|
||||
load_cmd = """{ echo 'SET FOREIGN_KEY_CHECKS=0;'; docker exec romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" -N -e "select concat(\\"DROP TABLE IF EXISTS \\`\\",table_name,\\"\\`;\\") from information_schema.tables where table_schema=\\"romm\\" and table_type=\\"BASE TABLE\\""' ; cat "$D"; } | docker exec -i romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" romm' > /root/dload.out 2>&1"""
|
||||
marker = "-- Dump completed"
|
||||
script = f"""
|
||||
D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql | head -1); echo "undo copy: $(basename $D) $(stat -c %s $D) B"
|
||||
docker stop {app} >/dev/null 2>&1; docker update --restart=no {app} >/dev/null
|
||||
t0=$(date +%s.%N)
|
||||
if tail -n 5 "$D" | grep -q -- '{marker}'; then echo "marker present"; else echo "MARKER ABSENT - refusing"; exit 0; fi
|
||||
{load_cmd}; echo "load rc=$?"; tail -2 /root/dload.out
|
||||
python3 -c "import time;print('marker check + load %.2fs' % (time.time()-$t0))"
|
||||
docker update --restart=unless-stopped {app} >/dev/null; docker start {app} >/dev/null
|
||||
"""
|
||||
w.say(w.guest(script, timeout=900))
|
||||
t0 = time.time()
|
||||
up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat"}[app], want=("200",), tries=60, delay=2)
|
||||
th = round(time.time() - t0, 1)
|
||||
A = FX[app].verify(w, s["sub"], s["seedA"], w.say)
|
||||
B = verify_b(app, s["sub"], s["seedA"], s["seedB"])
|
||||
st = db_state(app)
|
||||
w.say(f"D RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})")
|
||||
s["undoD"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st}
|
||||
save(app, s)
|
||||
|
||||
|
||||
def stage_cutoff(app):
|
||||
"""The cut-off copy, both methods, DETECTED before anything is loaded or swapped.
|
||||
F: a copy killed half-way (timeout) — compared with the source and checked for the finished-marker
|
||||
the build would write only after cp exits 0. D: the safety dump cut in half — the marker check."""
|
||||
s = load(app)
|
||||
big = max(s["volumes"], key=lambda v: int(w.guest(f"docker run --rm -v {v}:/v:ro alpine:3.20 du -sb /v | cut -f1").strip() or 0))
|
||||
script = f"""
|
||||
v={big}
|
||||
docker volume rm -f $v.cut >/dev/null 2>&1; docker volume create $v.cut >/dev/null
|
||||
cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'); docker stop $cs >/dev/null
|
||||
# Kill the copy CONTAINER, not the client (killing `docker run` leaves the container copying — measured
|
||||
# on docmost 2026-09-23, rc 137 with a complete copy and a marker).
|
||||
docker run -d --name cutcopy -v $v:/from:ro -v $v.cut:/to alpine:3.20 sh -c 'cp -a /from/. /to/ && touch /to/.felhom-copy-complete' >/dev/null
|
||||
sleep 0.02; docker kill cutcopy >/dev/null 2>&1; echo "copy container exit=$(docker wait cutcopy) (137 = killed)"; docker rm -f cutcopy >/dev/null 2>&1
|
||||
echo "source bytes=$(docker run --rm -v $v:/f:ro alpine:3.20 du -sb /f | cut -f1) cut copy bytes=$(docker run --rm -v $v.cut:/f:ro alpine:3.20 du -sb /f | cut -f1)"
|
||||
echo "finished-marker in the cut copy: $(docker run --rm -v $v.cut:/f:ro alpine:3.20 sh -c 'ls /f/.felhom-copy-complete 2>/dev/null | wc -l') -> F refuses to swap"
|
||||
docker volume rm -f $v.cut >/dev/null; docker start $cs >/dev/null
|
||||
D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql 2>/dev/null | head -1)
|
||||
if [ -n "$D" ]; then head -c $(( $(stat -c %s $D) / 2 )) $D > /root/cut.sql
|
||||
echo "D: whole copy marker: $(tail -n5 $D | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') cut copy marker: $(tail -n5 /root/cut.sql | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') -> D refuses to load"
|
||||
rm -f /root/cut.sql
|
||||
else echo "D: no safety dump exists for this app (R-641)"; fi
|
||||
"""
|
||||
w.say(f"=== cut-off copy: {app} (largest volume {big})")
|
||||
out = w.guest(script, timeout=600)
|
||||
w.say(out)
|
||||
s["cutoff"] = out
|
||||
save(app, s)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
w.login()
|
||||
{"prep": stage_prep, "fcopy": stage_fcopy, "break": stage_break, "update": stage_update,
|
||||
"undoF": stage_undoF, "undoD": stage_undoD, "cutoff": stage_cutoff,
|
||||
"state": lambda a: w.say(db_state(a))}[sys.argv[1]](sys.argv[2])
|
||||
@@ -0,0 +1,13 @@
|
||||
09:53:18 === prep docmost: deploy at the drill FROM pin
|
||||
09:53:18 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
|
||||
09:54:58 [1] deployed, controller state=running, pinned={'docmost': 'docmost/docmost:0.95.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'}
|
||||
09:54:58 deployed: True
|
||||
09:54:59 docmost: /api/auth/setup http=200 rc=0
|
||||
09:54:59 docmost: login as the seeded user http=200 ok=True
|
||||
09:54:59 C1 A: True
|
||||
09:54:59 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'}
|
||||
09:55:40 [4] backup idle; last=None
|
||||
09:55:40 docmost seed B: /api/spaces/create http=200 {"data":{"id":"01a0cd43-623a-7ddf-9b62-213653bc5558","name":"drillB3aab00","description":"","slug":"drillb3aab00","logo":null,"visibility":"private","defaultRol
|
||||
09:55:40 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False)
|
||||
09:55:40 B reads back: True
|
||||
09:55:43 named volumes: ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage']
|
||||
@@ -0,0 +1,8 @@
|
||||
09:55:53 db state before the update: 42 48 ledger, newest 20260620T010047-personal-spaces
|
||||
09:56:05 VOL docmost_docmost_postgres_data bytes=51156449 files=1561 | copy bytes=51156449 files=1561 | cp -a 0.98s | tar 0.54s (51838464 B)
|
||||
VOL docmost_docmost_redis_data bytes=39819 files=6 | copy bytes=39819 files=6 | cp -a 0.46s | tar 0.45s (37376 B)
|
||||
VOL docmost_docmost_storage bytes=4096 files=1 | copy bytes=4096 files=1 | cp -a 0.46s | tar 0.39s (1536 B)
|
||||
stop 0.52s copy(all, cp -a + the tar measurement) 9.00s start 0.75s
|
||||
free on the docker root: 21.04 GiB
|
||||
|
||||
09:56:17 front door answering again 12.2s after the start
|
||||
@@ -0,0 +1,9 @@
|
||||
09:56:25 [5] drill commit 386a34da911a: docmost docmost/docmost:0.95.0 -> docmost/docmost:0.96.0 (push rc=0)
|
||||
09:56:29 badge caught up after 4.5 s
|
||||
09:56:30 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
|
||||
09:56:30 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None
|
||||
09:56:31 + 1.0s phase=pulling label=Új verzió letöltése… err=None hold=None
|
||||
09:57:44 + 74.9s phase=starting label=Indítás az új verzióval… err=None hold=None
|
||||
09:57:46 + 76.9s phase=verifying label=Működés ellenőrzése… err=None hold=None
|
||||
09:59:18 + 168.1s phase=failed label=A frissítés nem sikerült err=A(z) docmost frissítése 2026-09-23 09:59-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 09:55 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza. hold=A(z) docmost frissítése 2026-09-23 09:59-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 09:55 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.
|
||||
09:59:18 {"final_phase": "failed", "hold_reason": "A(z) docmost frissítése 2026-09-23 09:59-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 09:55 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.", "state": "stopped", "duration_s": 168.1}
|
||||
@@ -0,0 +1,15 @@
|
||||
09:59:23 === undo by F: docmost
|
||||
09:59:41 [INFO] [settings] restore hold CLEARED for docmost
|
||||
|
||||
09:59:43 pinned_images:
|
||||
docmost: docmost/docmost:0.95.0
|
||||
docmost-postgres: postgres:16-alpine
|
||||
docmost-redis: redis:7-alpine
|
||||
|
||||
09:59:43 hold lift + pin back took 20.2 s (not part of a product undo)
|
||||
09:59:48 volumes put back in 2.20s
|
||||
|
||||
09:59:59 product start -> 200
|
||||
10:00:12 docmost: login as the seeded user http=200 ok=True
|
||||
10:00:12 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False)
|
||||
10:00:15 F RESULT: healthy=True after 23.6s A=True B=True db now: 48 52 ledger, newest 20260904T171920-public-spaces (before the update: 42 48 ledger, newest 20260620T010047-personal-spaces)
|
||||
@@ -0,0 +1,17 @@
|
||||
10:00:57 === REDO of the F undo with the bake-off's OWN pre-update definition (the first try read the re-captured unit — see docmost-40)
|
||||
10:01:00 image: docmost/docmost:0.95.0
|
||||
|
||||
10:01:02 pinned_images:
|
||||
docmost: docmost/docmost:0.95.0
|
||||
docmost-postgres: postgres:16-alpine
|
||||
docmost-redis: redis:7-alpine
|
||||
|
||||
10:01:07 stop + volumes put back in 2.91s
|
||||
|
||||
10:01:18 product start -> 200
|
||||
10:01:33 docmost/docmost:0.95.0 Up 15 seconds (healthy)
|
||||
successfully started
|
||||
|
||||
10:01:34 docmost: login as the seeded user http=200 ok=True
|
||||
10:01:34 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False)
|
||||
10:01:36 F RESULT (redo): healthy=True after 23.6s A=True B=True db now: 42 48 ledger, newest 20260620T010047-personal-spaces (before the update: 42 48 ledger, newest 20260620T010047-personal-spaces)
|
||||
@@ -0,0 +1,26 @@
|
||||
10:01:43 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
|
||||
10:01:43 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None
|
||||
10:01:44 + 1.1s phase=pulling label=Új verzió letöltése… err=None hold=None
|
||||
10:01:45 + 2.1s phase=starting label=Indítás az új verzióval… err=None hold=None
|
||||
10:01:46 + 3.1s phase=verifying label=Működés ellenőrzése… err=None hold=None
|
||||
10:03:18 + 95.4s phase=failed label=A frissítés nem sikerült err=A(z) docmost frissítése 2026-09-23 10:03-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 09:59 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza. hold=A(z) docmost frissítése 2026-09-23 10:03-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 09:59 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.
|
||||
10:03:18 {"final_phase": "failed", "hold_reason": "A(z) docmost frissítése 2026-09-23 10:03-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 09:59 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.", "state": "stopped", "duration_s": 95.4}
|
||||
10:03:18 === undo by D: docmost
|
||||
10:03:36 [INFO] [settings] restore hold CLEARED for docmost
|
||||
|
||||
10:03:38 pinned_images:
|
||||
docmost: docmost/docmost:0.95.0
|
||||
docmost-postgres: postgres:16-alpine
|
||||
docmost-redis: redis:7-alpine
|
||||
|
||||
10:03:38 hold lift + pin back took 20.2 s (not part of a product undo)
|
||||
10:04:01 undo copy: pre-restore-20260923T080143Z-docmost-postgres.sql 136043 B
|
||||
marker present
|
||||
load rc=0
|
||||
ALTER TABLE
|
||||
ALTER TABLE
|
||||
marker check + load 1.87s
|
||||
|
||||
10:04:13 docmost: login as the seeded user http=200 ok=True
|
||||
10:04:14 docmost B: /api/spaces http=200 seeded-space-listed=True (negative control listed=False)
|
||||
10:04:16 D RESULT: healthy=True after 12.2s A=True B=True db now: 42 48 ledger, newest 20260620T010047-personal-spaces (before the update: 42 48 ledger, newest 20260620T010047-personal-spaces)
|
||||
@@ -0,0 +1,16 @@
|
||||
10:04:32 === cut-off copy: docmost (largest volume docmost_docmost_postgres_data)
|
||||
10:04:39 copy rc=137 (137 = killed)
|
||||
source bytes=69334709 cut copy bytes=69334709
|
||||
finished-marker in the cut copy: 1 -> F refuses to swap
|
||||
D: whole copy marker: 1 cut copy marker: 0 -> D refuses to load
|
||||
|
||||
10:04:55 NOTE: the attempt above killed the docker CLIENT (rc 137) while the copy CONTAINER ran on and finished — so nothing was cut and the marker was written. Redo: kill the copy CONTAINER itself mid-copy.
|
||||
10:05:07 kill after 0.15s: container exit=0 source bytes=69172397 cut copy bytes=69172397 finished-marker=1
|
||||
kill after 0.25s: container exit=0 source bytes=69172397 cut copy bytes=69172397 finished-marker=1
|
||||
kill after 0.35s: container exit=0 source bytes=69172397 cut copy bytes=69172397 finished-marker=1
|
||||
|
||||
10:05:31 kill after 0s: container exit=137 source bytes=69172397 cut copy bytes=42921516 finished-marker=0
|
||||
kill after 0.02s: container exit=137 source bytes=69172397 cut copy bytes=48586287 finished-marker=0
|
||||
kill after 0.05s: container exit=137 source bytes=69172397 cut copy bytes=59159114 finished-marker=0
|
||||
kill after 0.08s: container exit=137 source bytes=69172397 cut copy bytes=68445276 finished-marker=0
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,60 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Point guest 9202 at the drill catalog (and a 90 s health timeout), or restore the saved config.
|
||||
|
||||
`09` §6.5: `git.repo_url` alone is INERT (R-615) — the cache dir must go too. The saved copy is
|
||||
`controller.yaml.pre-bakeoff` (NOT the older `.pre-28`, which a restore must never pick up).
|
||||
"""
|
||||
import re, sys, io
|
||||
sys.path.insert(0, '.')
|
||||
import walk as w
|
||||
|
||||
VOL = "/var/lib/docker/volumes/felhom-controller-data/_data"
|
||||
DRILL_REPO = "https://gitea.dooplex.hu/admin/app-catalog-drill.git"
|
||||
|
||||
|
||||
def creds():
|
||||
for l in io.open("/home/kisfenyo/.git-credentials").read().strip().split("\n"):
|
||||
m = re.match(r'https://(admin):([^@]+)@gitea\.dooplex\.hu', l)
|
||||
if m:
|
||||
return m.group(1), m.group(2)
|
||||
raise SystemExit("no admin credential")
|
||||
|
||||
|
||||
def to_drill():
|
||||
u, t = creds()
|
||||
print(w.guest(f"""
|
||||
set -e
|
||||
test -f {VOL}/controller.yaml.pre-bakeoff || cp -p {VOL}/controller.yaml {VOL}/controller.yaml.pre-bakeoff
|
||||
python3 - <<'PY'
|
||||
import re
|
||||
p = "{VOL}/controller.yaml"
|
||||
s = open(p).read()
|
||||
s = re.sub(r'(^\\s+repo_url: ).*$', r'\\g<1>{DRILL_REPO}', s, count=1, flags=re.M)
|
||||
s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+token: ).*$', r'\\g<1>"{t}"', s, count=1, flags=re.M)
|
||||
s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+username: ).*$', r'\\g<1>"{u}"', s, count=1, flags=re.M)
|
||||
if not re.search(r'^update:', s, re.M):
|
||||
s += "update:\\n health_timeout: 90s\\n"
|
||||
open(p, "w").write(s)
|
||||
PY
|
||||
rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache
|
||||
docker restart felhom-controller >/dev/null
|
||||
sleep 15
|
||||
grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: <redacted>/'
|
||||
grep -A2 '^update:' {VOL}/controller.yaml
|
||||
"""))
|
||||
|
||||
|
||||
def restore():
|
||||
print(w.guest(f"""
|
||||
set -e
|
||||
cp -p {VOL}/controller.yaml.pre-bakeoff {VOL}/controller.yaml
|
||||
rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache
|
||||
docker restart felhom-controller >/dev/null
|
||||
sleep 15
|
||||
grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: <redacted>/'
|
||||
grep -c '^update:' {VOL}/controller.yaml || true
|
||||
"""))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
to_drill() if sys.argv[1] == "drill" else restore()
|
||||
@@ -0,0 +1,15 @@
|
||||
10:05:44 === prep romm: deploy at the drill FROM pin
|
||||
10:05:47 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/romm']
|
||||
10:05:47 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['HDD_PATH']
|
||||
10:05:47 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
|
||||
10:07:02 [1] deployed, controller state=running, pinned={'romm': 'rommapp/romm:5.0.0', 'romm-db': 'mariadb:11.4', 'romm-redis': 'redis:7-alpine'}
|
||||
10:07:02 deployed: True
|
||||
10:07:12 romm: POST /api/users http=201
|
||||
10:07:13 romm: login as the seeded user http=200 ok=True
|
||||
10:07:13 C1 A: True
|
||||
10:07:13 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'}
|
||||
10:08:24 [4] backup idle; last=None
|
||||
10:08:54 romm seed B: POST /api/users (as the admin) http=201 {"id":2,"username":"drillb12740b","email":"drillb12740b@gate.invalid","enabled":true,"role":"user","permission_group_id"
|
||||
10:08:54 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False)
|
||||
10:08:54 B reads back: True
|
||||
10:08:57 named volumes: ['romm_romm_config', 'romm_romm_db_data', 'romm_romm_redis_data']
|
||||
@@ -0,0 +1,10 @@
|
||||
10:08:59 db state before the update: tables=TABLE" romm: 1: Syntax error: Unterminated quoted string alembic=0095_virtual_collections_source
|
||||
10:09:01 image: rommapp/romm:5.0.0
|
||||
|
||||
10:09:15 VOL romm_romm_config bytes=4158 files=2 | copy bytes=4158 files=2 | cp -a 0.40s | tar 0.38s (2560 B)
|
||||
VOL romm_romm_db_data bytes=160302192 files=273 | copy bytes=160302192 files=273 | cp -a 0.68s | tar 0.50s (160458240 B)
|
||||
VOL romm_romm_redis_data bytes=14512195 files=6 | copy bytes=14512195 files=6 | cp -a 0.42s | tar 0.40s (14509056 B)
|
||||
stop 3.79s copy(all, cp -a + the tar measurement) 7.75s start 0.68s
|
||||
free on the docker root: 19.04 GiB
|
||||
|
||||
10:09:49 front door answering again 33.8s after the start
|
||||
@@ -0,0 +1,2 @@
|
||||
10:09:50 [5] drill commit de819653d3d0: romm rommapp/romm:5.0.0 -> rommapp/romm:5.3.0 (push rc=0)
|
||||
10:09:55 badge caught up after 4.4 s
|
||||
@@ -0,0 +1,7 @@
|
||||
10:09:55 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
|
||||
10:09:55 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None
|
||||
10:09:56 + 1.0s phase=pulling label=Új verzió letöltése… err=None hold=None
|
||||
10:10:08 + 13.3s phase=starting label=Indítás az új verzióval… err=None hold=None
|
||||
10:10:14 + 19.5s phase=verifying label=Működés ellenőrzése… err=None hold=None
|
||||
10:11:54 + 118.9s phase=failed label=A frissítés nem sikerült err=A(z) romm frissítése 2026-09-23 10:11-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:08 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem. hold=A(z) romm frissítése 2026-09-23 10:11-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:08 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.
|
||||
10:11:54 {"final_phase": "failed", "hold_reason": "A(z) romm frissítése 2026-09-23 10:11-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:08 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.", "state": "stopped", "duration_s": 119.0}
|
||||
@@ -0,0 +1,15 @@
|
||||
10:11:54 === undo by F: romm
|
||||
10:12:12 [INFO] [settings] restore hold CLEARED for romm
|
||||
|
||||
10:12:14 pinned_images:
|
||||
romm: rommapp/romm:5.0.0
|
||||
romm-db: mariadb:11.4
|
||||
romm-redis: redis:7-alpine
|
||||
|
||||
10:12:14 hold lift + pin back took 20.2 s (not part of a product undo)
|
||||
10:12:18 volumes put back in 1.64s
|
||||
|
||||
10:12:30 product start -> 200
|
||||
10:13:05 romm: login as the seeded user http=200 ok=True
|
||||
10:13:05 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False)
|
||||
10:13:08 F RESULT: healthy=True after 45.3s A=True B=True db now: tables=TABLE" romm: 1: Syntax error: Unterminated quoted string alembic=0095_virtual_collections_source (before the update: tables=TABLE" romm: 1: Syntax error: Unterminated quoted string alembic=0095_virtual_collections_source)
|
||||
@@ -0,0 +1,7 @@
|
||||
10:13:09 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
|
||||
10:13:09 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None
|
||||
10:13:10 + 1.1s phase=pulling label=Új verzió letöltése… err=None hold=None
|
||||
10:13:11 + 2.1s phase=starting label=Indítás az új verzióval… err=None hold=None
|
||||
10:13:14 + 5.2s phase=verifying label=Működés ellenőrzése… err=None hold=None
|
||||
10:14:48 + 99.5s phase=failed label=A frissítés nem sikerült err=A(z) romm frissítése 2026-09-23 10:14-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:11 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem. hold=A(z) romm frissítése 2026-09-23 10:14-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:11 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.
|
||||
10:14:48 {"final_phase": "failed", "hold_reason": "A(z) romm frissítése 2026-09-23 10:14-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:11 — ez a másolat csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem.", "state": "stopped", "duration_s": 99.5}
|
||||
@@ -0,0 +1,17 @@
|
||||
10:14:48 === undo by D: romm
|
||||
10:15:06 [INFO] [settings] restore hold CLEARED for romm
|
||||
|
||||
10:15:08 pinned_images:
|
||||
romm: rommapp/romm:5.0.0
|
||||
romm-db: mariadb:11.4
|
||||
romm-redis: redis:7-alpine
|
||||
|
||||
10:15:08 hold lift + pin back took 20.1 s (not part of a product undo)
|
||||
10:15:40 undo copy: pre-restore-20260923T081309Z-romm-mariadb.sql 62943 B
|
||||
marker present
|
||||
load rc=0
|
||||
marker check + load 1.41s
|
||||
|
||||
10:16:15 romm: login as the seeded user http=200 ok=True
|
||||
10:16:16 romm B: GET /api/users http=200 seeded-user-listed=True (negative control listed=False)
|
||||
10:16:18 D RESULT: healthy=True after 33.6s A=True B=True db now: tables=24 alembic=0095_virtual_collections_source (before the update: tables=TABLE" romm: 1: Syntax error: Unterminated quoted string alembic=0095_virtual_collections_source)
|
||||
@@ -0,0 +1,6 @@
|
||||
10:16:26 === cut-off copy: romm (largest volume romm_romm_db_data)
|
||||
10:16:35 copy container exit=137 (137 = killed)
|
||||
source bytes=160215912 cut copy bytes=105844788
|
||||
finished-marker in the cut copy: 0 -> F refuses to swap
|
||||
D: whole copy marker: 1 cut copy marker: 0 -> D refuses to load
|
||||
|
||||
@@ -0,0 +1,160 @@
|
||||
#!/usr/bin/env python3
|
||||
"""spike.py — Part 1 of the 2026-09-23 brief: the AUTOMATIC UNDO, performed BY HAND on guest 9202.
|
||||
|
||||
EVIDENCE, NOT PRODUCT. Every product act goes through the endpoints the UI invokes (deploy, backup,
|
||||
sync, rescan, update, remove). The UNDO itself has no product path yet (that is what is being
|
||||
spiked), so it is performed by hand with plain docker/compose inside the guest, using exactly the
|
||||
steps the product would take, each one timed.
|
||||
|
||||
State between stages lives in state-<app>.json so each stage can be run, read, and only then
|
||||
followed by the next (the hand undo needs a person looking at what the previous step left).
|
||||
"""
|
||||
import json, os, sys, time, re
|
||||
sys.path.insert(0, ".")
|
||||
import walk as w
|
||||
import fixtures as fx
|
||||
|
||||
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
FX = {"docmost": fx.Docmost(), "vikunja": fx.Vikunja(), "romm": fx.Romm()}
|
||||
SUB = {"docmost": "docs", "vikunja": "tasks", "romm": "arcade"}
|
||||
|
||||
|
||||
def st_path(app):
|
||||
return os.path.join(HERE, f"state-{app}.json")
|
||||
|
||||
|
||||
def load(app):
|
||||
return json.load(open(st_path(app))) if os.path.exists(st_path(app)) else {}
|
||||
|
||||
|
||||
def save(app, s):
|
||||
json.dump(s, open(st_path(app), "w"), indent=2, ensure_ascii=False)
|
||||
|
||||
|
||||
def ts():
|
||||
return time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
|
||||
|
||||
|
||||
# ---- a SECOND seed, written AFTER the backup and BEFORE the update. It is the discriminator: only
|
||||
# the pre-pin safety dump can hold it — the backup tier copy was taken before it existed. So if it
|
||||
# reads back after the undo, the undo used the safety dump; if only A reads back, it used the tier.
|
||||
def docmost_seed_b(sub, A):
|
||||
jar = "/tmp/dm.jar"
|
||||
w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json",
|
||||
data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST")
|
||||
name = "drillB" + os.urandom(3).hex()
|
||||
rc, code, out = w.app_curl(sub, "/api/spaces/create", "-b", jar, "-H", "Content-Type: application/json",
|
||||
data=json.dumps({"name": name, "slug": name.lower()}), method="POST")
|
||||
w.say(f" docmost seed B: /api/spaces/create http={code} {out[:160]}")
|
||||
return {"space": name} if code in ("200", "201") else None
|
||||
|
||||
|
||||
def docmost_verify_b(sub, A, B):
|
||||
jar = "/tmp/dm.jar"
|
||||
rc, code, out = w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json",
|
||||
data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST")
|
||||
if code not in ("200", "201"):
|
||||
w.say(f" docmost B: cannot log in (http={code})"); return False
|
||||
rc, code, out = w.app_curl(sub, "/api/spaces", "-b", jar, "-H", "Content-Type: application/json",
|
||||
data="{}", method="POST")
|
||||
ok = code in ("200", "201") and B["space"] in out
|
||||
neg = "drillBnever" in out
|
||||
w.say(f" docmost B: /api/spaces http={code} seeded-space-listed={ok} (negative control listed={neg})")
|
||||
return ok and not neg
|
||||
|
||||
|
||||
def break_edge(app, frm, to, port_from, port_to):
|
||||
"""The failing edge: a REAL migrating image move, plus — in the DRILL template only — the
|
||||
health probe pointed at a port the app does not answer. Both in one drill commit."""
|
||||
fy = f"{w.DRILL}/templates/{app}/.felhom.yml"
|
||||
f = open(fy).read()
|
||||
m = re.search(r"(healthcheck:\n(?:.*\n){0,8}?\s+port: )" + str(port_from) + r"\b", f)
|
||||
assert m, "probe port not found"
|
||||
f = f[:m.end() - len(str(port_from))] + str(port_to) + f[m.end():]
|
||||
open(fy, "w").write(f)
|
||||
h = w.drill_bump(app, frm, to)
|
||||
return h
|
||||
|
||||
|
||||
def pg_state(container, db, user):
|
||||
return w.guest(f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1; "
|
||||
f"docker exec {container} psql -U {user} -d {db} -Atc \"select name from kysely_migration order by name desc limit 3\" 2>&1; "
|
||||
f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from kysely_migration\" 2>&1")
|
||||
|
||||
|
||||
def romm_seed_b(sub, A):
|
||||
"""A SECOND RomM user, created by the first (admin) one — written after the backup."""
|
||||
jar, tok = FX["romm"]._csrf(w, sub)
|
||||
user = "drillb" + os.urandom(3).hex()
|
||||
rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}",
|
||||
"-u", f"{A['user']}:{A['pw']}", "-H", "Content-Type: application/json",
|
||||
data=json.dumps({"username": user, "email": f"{user}@gate.invalid",
|
||||
"password": "Drill-" + os.urandom(8).hex(), "role": "viewer"}),
|
||||
method="POST")
|
||||
w.say(f" romm seed B: POST /api/users (as the admin) http={code} {body[:120]}")
|
||||
return {"user": user} if code in ("200", "201") else None
|
||||
|
||||
|
||||
def romm_verify_b(sub, A, B):
|
||||
jar, tok = FX["romm"]._csrf(w, sub)
|
||||
rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}",
|
||||
"-u", f"{A['user']}:{A['pw']}")
|
||||
ok = code == "200" and B["user"] in body
|
||||
neg = "drillbnever" in body
|
||||
w.say(f" romm B: GET /api/users http={code} seeded-user-listed={ok} (negative control listed={neg})")
|
||||
return ok and not neg
|
||||
|
||||
|
||||
def tree(paths):
|
||||
cmd = "; ".join(f"echo \"{p}: files=$(find {p} -type f 2>/dev/null | wc -l) sum=$(find {p} -type f -exec sha256sum {{}} + 2>/dev/null | sort | sha256sum | cut -c1-16)\"" for p in paths)
|
||||
return w.guest(cmd)
|
||||
|
||||
|
||||
def my_state(container="romm-db", db="romm"):
|
||||
# The root password is used INSIDE the container from its own env — it never leaves it.
|
||||
q = lambda sql: f"docker exec {container} sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" {db}' 2>&1 | tr '\\n' ' '"
|
||||
return w.guest(f"""echo -n "tables=$({q("select count(*) from information_schema.tables where table_schema=database()")}) "
|
||||
echo -n "alembic=$({q("select version_num from alembic_version")}) "
|
||||
echo -n "users=$({q("select count(*) from users")})"
|
||||
""")
|
||||
|
||||
|
||||
def vik_seed_b(sub, A):
|
||||
"""A second project, created AFTER the backup — plus a task with a real ATTACHMENT (a file the
|
||||
app writes into its files volume), uploaded through the app's own attachment API."""
|
||||
tok, why = FX["vikunja"]._token(w, sub, A)
|
||||
title = "drillB-" + os.urandom(4).hex()
|
||||
rc, code, body = w.app_curl(sub, "/api/v1/projects", "-H", f"Authorization: Bearer {tok}",
|
||||
"-H", "Content-Type: application/json", data=json.dumps({"title": title}), method="PUT")
|
||||
if code not in ("200", "201"):
|
||||
w.say(f" vikunja B: project refused {code} {body[:120]}"); return None
|
||||
pid = json.loads(body)["id"]
|
||||
rc, code, body = w.app_curl(sub, f"/api/v1/projects/{pid}/tasks", "-H", f"Authorization: Bearer {tok}",
|
||||
"-H", "Content-Type: application/json", data=json.dumps({"title": "task-" + title}), method="PUT")
|
||||
tid = json.loads(body)["id"] if code in ("200", "201") else None
|
||||
content = "drill attachment " + os.urandom(8).hex()
|
||||
fn = "/tmp/vik-att.txt"; open(fn, "w").write(content)
|
||||
rc, code2, body2 = w.app_curl(sub, f"/api/v1/tasks/{tid}/attachments", "-H", f"Authorization: Bearer {tok}",
|
||||
"-F", f"files=@{fn}", method="PUT")
|
||||
w.say(f" vikunja seed B: project http=200 task={tid} attachment upload http={code2} {body2[:120]}")
|
||||
return {"title": title, "pid": pid, "tid": tid, "att": content}
|
||||
|
||||
|
||||
def vik_verify_b(sub, A, B):
|
||||
tok, why = FX["vikunja"]._token(w, sub, A)
|
||||
if not tok:
|
||||
w.say(f" vikunja B: cannot log in {why}"); return {"project": False, "attachment": False}
|
||||
rc, code, body = w.app_curl(sub, f"/api/v1/projects/{B['pid']}", "-H", f"Authorization: Bearer {tok}")
|
||||
proj = code == "200" and B["title"] in body
|
||||
rc, code, body = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments", "-H", f"Authorization: Bearer {tok}")
|
||||
att_ok = False
|
||||
try:
|
||||
atts = json.loads(body)
|
||||
if atts:
|
||||
aid = atts[0]["id"]
|
||||
rc, c3, b3 = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments/{aid}", "-H", f"Authorization: Bearer {tok}")
|
||||
att_ok = c3 == "200" and B["att"] in b3
|
||||
except Exception as e:
|
||||
w.say(f" vikunja B: attachments list unreadable http={code} {body[:120]}")
|
||||
w.say(f" vikunja B: project readback={proj} attachment content readback={att_ok}")
|
||||
return {"project": proj, "attachment": att_ok}
|
||||
@@ -0,0 +1,14 @@
|
||||
10:16:50 === prep vikunja: deploy at the drill FROM pin
|
||||
10:16:50 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
|
||||
10:16:55 [1] deployed, controller state=running, pinned={'vikunja': 'vikunja/vikunja:2.3.0'}
|
||||
10:16:55 deployed: True
|
||||
10:16:56 vikunja: register http=200
|
||||
10:16:56 vikunja: create project http=201
|
||||
10:16:56 vikunja: readback of the seeded project http=200 ok=True
|
||||
10:16:56 C1 A: True
|
||||
10:16:56 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'}
|
||||
10:17:56 [4] backup idle; last=None
|
||||
10:17:57 vikunja seed B: project http=200 task=1 attachment upload http=200 {"errors":null,"success":[{"id":1,"task_id":1,"created_by":{"id":1,"name":"","username":"drill6d84b2","created":"2026-09
|
||||
10:17:57 vikunja B: project readback=True attachment content readback=True
|
||||
10:17:57 B reads back: True
|
||||
10:18:00 named volumes: ['vikunja_vikunja_data', 'vikunja_vikunja_db']
|
||||
@@ -0,0 +1,9 @@
|
||||
10:18:03 db state before the update: tables=36 ledger=117 newest=SCHEMA_INIT
|
||||
10:18:05 image: vikunja/vikunja:2.3.0
|
||||
|
||||
10:18:12 VOL vikunja_vikunja_data bytes=4129 files=2 | copy bytes=4129 files=2 | cp -a 0.42s | tar 0.41s (2560 B)
|
||||
VOL vikunja_vikunja_db bytes=2471792 files=4 | copy bytes=2471792 files=4 | cp -a 0.42s | tar 0.40s (2470912 B)
|
||||
stop 0.18s copy(all, cp -a + the tar measurement) 5.03s start 0.23s
|
||||
free on the docker root: 18.18 GiB
|
||||
|
||||
10:18:14 front door answering again 2.1s after the start
|
||||
@@ -0,0 +1,2 @@
|
||||
10:18:15 [5] drill commit 9772e7c0b9f5: vikunja vikunja/vikunja:2.3.0 -> vikunja/vikunja:2.6.0 (push rc=0)
|
||||
10:18:20 badge caught up after 4.5 s
|
||||
@@ -0,0 +1,7 @@
|
||||
10:18:20 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
|
||||
10:18:20 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None
|
||||
10:18:21 + 1.0s phase=pulling label=Új verzió letöltése… err=None hold=None
|
||||
10:18:24 + 4.1s phase=starting label=Indítás az új verzióval… err=None hold=None
|
||||
10:18:25 + 5.1s phase=verifying label=Működés ellenőrzése… err=None hold=None
|
||||
10:19:56 + 95.4s phase=failed label=A frissítés nem sikerült err=A(z) vikunja frissítése 2026-09-23 10:19-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:17 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza. hold=A(z) vikunja frissítése 2026-09-23 10:19-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:17 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.
|
||||
10:19:56 {"final_phase": "failed", "hold_reason": "A(z) vikunja frissítése 2026-09-23 10:19-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:17 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.", "state": "stopped", "duration_s": 95.4}
|
||||
@@ -0,0 +1,13 @@
|
||||
10:19:56 === undo by F: vikunja
|
||||
10:20:14 [INFO] [settings] restore hold CLEARED for vikunja
|
||||
|
||||
10:20:16 pinned_images:
|
||||
vikunja: vikunja/vikunja:2.3.0
|
||||
|
||||
10:20:16 hold lift + pin back took 20.2 s (not part of a product undo)
|
||||
10:20:19 volumes put back in 0.90s
|
||||
|
||||
10:20:19 product start -> 200
|
||||
10:20:22 vikunja: readback of the seeded project http=200 ok=True
|
||||
10:20:22 vikunja B: project readback=True attachment content readback=True
|
||||
10:20:24 F RESULT: healthy=True after 2.7s A=True B=True db now: tables=36 ledger=117 newest=SCHEMA_INIT (before the update: tables=36 ledger=117 newest=SCHEMA_INIT)
|
||||
@@ -0,0 +1,6 @@
|
||||
10:20:25 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
|
||||
10:20:25 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None
|
||||
10:20:26 + 1.1s phase=pulling label=Új verzió letöltése… err=None hold=None
|
||||
10:20:27 + 2.1s phase=verifying label=Működés ellenőrzése… err=None hold=None
|
||||
10:21:58 + 93.3s phase=failed label=A frissítés nem sikerült err=A(z) vikunja frissítése 2026-09-23 10:21-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:19 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza. hold=A(z) vikunja frissítése 2026-09-23 10:21-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:19 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.
|
||||
10:21:58 {"final_phase": "failed", "hold_reason": "A(z) vikunja frissítése 2026-09-23 10:21-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-23 10:19 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.", "state": "stopped", "duration_s": 93.3}
|
||||
@@ -0,0 +1,8 @@
|
||||
10:21:58 === undo by D: vikunja
|
||||
10:22:16 [INFO] [settings] restore hold CLEARED for vikunja
|
||||
|
||||
10:22:18 pinned_images:
|
||||
vikunja: vikunja/vikunja:2.3.0
|
||||
|
||||
10:22:18 hold lift + pin back took 20.1 s (not part of a product undo)
|
||||
10:22:18 vikunja has no database server: the product wrote NO safety dump (R-641). D's copy for it is a volume tar at safety-dump time, which is F with extra steps — recorded, not re-run.
|
||||
@@ -0,0 +1,11 @@
|
||||
10:22:23 === cut-off copy: vikunja (largest volume vikunja_vikunja_db)
|
||||
10:22:27 copy container exit=0 (137 = killed)
|
||||
source bytes=2920872 cut copy bytes=2920872
|
||||
finished-marker in the cut copy: 1 -> F refuses to swap
|
||||
D: no safety dump exists for this app (R-641)
|
||||
|
||||
10:22:40 RETRY — 2.9 MB copied inside 0.02 s, so the first cut missed; kill with no delay:
|
||||
10:22:48 try 1: container exit=0 source bytes=2920872 cut copy bytes=2920872 finished-marker=1
|
||||
try 2: container exit=0 source bytes=2920872 cut copy bytes=2920872 finished-marker=1
|
||||
try 3: container exit=0 source bytes=2920872 cut copy bytes=2920872 finished-marker=1
|
||||
|
||||
@@ -0,0 +1,493 @@
|
||||
#!/usr/bin/env python3
|
||||
"""walk.py — ONE app's full update walk on guest 9202, through the product's own endpoints.
|
||||
|
||||
EVIDENCE, NOT PRODUCT. It presses exactly the buttons a person presses:
|
||||
POST /api/stacks/<n>/deploy · POST /api/backup/run · POST /api/sync · POST /api/stacks/rescan
|
||||
POST /api/stacks/<n>/update · POST /api/stacks/<n>/remove
|
||||
and reads GET /api/stacks/<n>. No controller code exists for it.
|
||||
|
||||
The walk, per `09` §6.4 and the update-night brief §4:
|
||||
1 deploy from the DRILL catalog at the LIVE pin
|
||||
2 seed through the app's OWN front door (R-156: never a volume, never SQL)
|
||||
3 read the seed back <- control C1; a fixture that cannot prove itself proves nothing
|
||||
4 „Mentés most"
|
||||
5 commit the real one-step bump to the DRILL repo, sync, rescan, read the badge in BOTH languages
|
||||
6 press the guarded Update, record every phase with timestamps
|
||||
7 read the seed back through the front door
|
||||
8 the four version observables side by side
|
||||
9 write the verdict record in `09`'s JSON shape
|
||||
|
||||
`inconclusive` is a first-class verdict and is NEVER collapsed into `failed`.
|
||||
"""
|
||||
import argparse, json, os, re, subprocess, sys, time
|
||||
from datetime import datetime, timezone
|
||||
|
||||
SC = "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/6e5a1a3b-6d8c-4ee1-bc3f-c555eb3f7578/scratchpad"
|
||||
EV = "/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/undo-bakeoff-2026-09-23"
|
||||
DRILL = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill"
|
||||
BASE = "https://192.168.0.114"
|
||||
HOSTHDR = "Host: felhom.enkisfelhom.hu"
|
||||
DOMAIN = "enkisfelhom.hu"
|
||||
HP = "demo-hp"
|
||||
|
||||
LOG = []
|
||||
|
||||
|
||||
def say(*a):
|
||||
line = " ".join(str(x) for x in a)
|
||||
ts = datetime.now().strftime("%H:%M:%S")
|
||||
print(f"{ts} {line}", flush=True)
|
||||
LOG.append(f"{ts} {line}")
|
||||
|
||||
|
||||
def sh(args, timeout=300, inp=None):
|
||||
try:
|
||||
return subprocess.run(args, capture_output=True, text=True, timeout=timeout, input=inp)
|
||||
except (subprocess.TimeoutExpired, OSError) as e:
|
||||
return subprocess.CompletedProcess(args, 124, "", f"{e}")
|
||||
|
||||
|
||||
def guest(script, timeout=600):
|
||||
"""Run a bash script inside guest 9202. Piped as a file — never as an argument (quoting)."""
|
||||
r = sh(["ssh", "-o", "ConnectTimeout=20", "-o", "StrictHostKeyChecking=accept-new", HP,
|
||||
"cat > /tmp/w.sh; pct push 9202 /tmp/w.sh /tmp/w.sh >/dev/null 2>&1; "
|
||||
"pct exec 9202 -- bash /tmp/w.sh; rm -f /tmp/w.sh"],
|
||||
timeout=timeout, inp=script)
|
||||
return r.stdout or ""
|
||||
|
||||
|
||||
def login():
|
||||
pw = open(f"{SC}/.ctlpw").read().strip()
|
||||
sh(["curl", "-sk", "-D", f"{SC}/hdr.txt", "-o", "/dev/null", "-H", HOSTHDR,
|
||||
"-X", "POST", "--data-urlencode", f"password={pw}", f"{BASE}/login"])
|
||||
h = open(f"{SC}/hdr.txt").read()
|
||||
m = re.search(r"felhom_session=[A-Za-z0-9._-]+", h, re.I)
|
||||
if not m:
|
||||
sys.exit("login failed: no session cookie")
|
||||
open(f"{SC}/sess.txt", "w").write(m.group(0))
|
||||
r = sh(["curl", "-sk", "-L", "-H", HOSTHDR, "-H", f"Cookie: {m.group(0)}", f"{BASE}/"])
|
||||
c = re.search(r'<meta name="csrf-token" content="([^"]+)"', r.stdout or "")
|
||||
if not c:
|
||||
sys.exit("login failed: no csrf token")
|
||||
open(f"{SC}/csrf.txt", "w").write(c.group(1))
|
||||
|
||||
|
||||
def ctl(method, path, data=None, raw=False, tries=2):
|
||||
"""One controller API call. Re-logs in once on a 302/401 — the controller's session store is
|
||||
in memory, so any controller restart during the night invalidates it silently."""
|
||||
for attempt in range(tries):
|
||||
sess = open(f"{SC}/sess.txt").read().strip()
|
||||
csrf = open(f"{SC}/csrf.txt").read().strip()
|
||||
args = ["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", "-w", "\n%{http_code}"]
|
||||
if method != "GET":
|
||||
args += ["-H", f"X-CSRF-Token: {csrf}", "-H", "Content-Type: application/json",
|
||||
"-X", method]
|
||||
if data is not None:
|
||||
args += ["--data", json.dumps(data)]
|
||||
args.append(f"{BASE}{path}")
|
||||
r = sh(args)
|
||||
body, _, code = (r.stdout or "").rpartition("\n")
|
||||
if code.strip() in ("302", "401") and attempt + 1 < tries:
|
||||
login()
|
||||
continue
|
||||
if raw:
|
||||
return code.strip(), body
|
||||
try:
|
||||
return code.strip(), json.loads(body)
|
||||
except Exception:
|
||||
return code.strip(), {"_raw": body[:600]}
|
||||
return code.strip(), {"_raw": body[:600]}
|
||||
|
||||
|
||||
def page(path):
|
||||
sess = open(f"{SC}/sess.txt").read().strip()
|
||||
r = sh(["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", f"{BASE}{path}"])
|
||||
return r.stdout or ""
|
||||
|
||||
|
||||
def app_curl(sub, path, *extra, method=None, data=None, timeout=45):
|
||||
"""A call to the APP's own front door on 9202 — the household's route, not ours."""
|
||||
args = ["curl", "-sSk", "--max-time", str(timeout), "-H", f"Host: {sub}.{DOMAIN}",
|
||||
"-w", "\n%{http_code}"]
|
||||
if method:
|
||||
args += ["-X", method]
|
||||
if data is not None:
|
||||
args += ["--data-binary", "@-"]
|
||||
args += list(extra) + [f"{BASE}{path}"]
|
||||
r = sh(args, timeout=timeout + 30, inp=data)
|
||||
body, _, code = (r.stdout or "").rpartition("\n")
|
||||
return r.returncode, code.strip(), body
|
||||
|
||||
|
||||
def stack(name):
|
||||
_, d = ctl("GET", f"/api/stacks/{name}")
|
||||
return (d.get("data") or {}) if isinstance(d, dict) else {}
|
||||
|
||||
|
||||
def wait_app(sub, path="/", want=("200", "302", "303", "401", "403"), tries=60, delay=5):
|
||||
"""Settling says the container runs; this says the APP answers. Not the same thing."""
|
||||
last = None
|
||||
for _ in range(tries):
|
||||
rc, code, _ = app_curl(sub, path, timeout=15)
|
||||
last = (rc, code)
|
||||
if rc == 0 and code in want:
|
||||
return True
|
||||
time.sleep(delay)
|
||||
say(f" app never answered on {sub}{path} (last rc={last[0]} code={last[1]})")
|
||||
return False
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ the walk
|
||||
|
||||
|
||||
DRIVE = "/mnt/felhom-drives/scratch_hdd/userdata"
|
||||
|
||||
# What THIS run generated for a deploy, per app. Deploy secrets are ENCRYPTED AT REST in
|
||||
# `app.yaml` (`ENC:…`), which is right and which means a fixture cannot read an app's admin
|
||||
# password back off the box — the household sees it once. So the value the harness itself
|
||||
# generated is kept here for the life of the run, and nowhere else.
|
||||
GENERATED = {}
|
||||
|
||||
|
||||
def deploy_values(name, sub):
|
||||
"""Fill EVERY required deploy field the way the wizard would, by asking the box what this app
|
||||
asks for — `GET /api/stacks/<n>/deploy-fields` — instead of assuming DOMAIN+SUBDOMAIN.
|
||||
|
||||
Measured 2026-09-21: three apps in one batch refused at the deploy with a correct 400 because
|
||||
a required field was absent — `HDD_PATH` (navidrome, audiobookshelf) and an admin password
|
||||
(grafana). The refusals happen BEFORE anything is created (`deploy.go:324`), which is the only
|
||||
reason this was safe to discover by running it (live-probes rule).
|
||||
|
||||
A `path` field must name a directory that ALREADY EXISTS (`deploy.go:330`), so one is made on
|
||||
the scratch drive first — the same act the drive browser performs for a household.
|
||||
"""
|
||||
code, d = ctl("GET", f"/api/stacks/{name}/deploy-fields")
|
||||
fields = (((d.get("data") or {}).get("metadata") or {}).get("deploy_fields")) or []
|
||||
values = {"DOMAIN": DOMAIN, "SUBDOMAIN": sub}
|
||||
made = []
|
||||
for f in fields:
|
||||
ev, ty = f.get("env_var"), f.get("type")
|
||||
if ev in values:
|
||||
continue
|
||||
# `type: password` is MANDATORY whatever `required` says — `deploy.go:305-312` refuses
|
||||
# when the caller sends none, deliberately ("the user needs to know their password"),
|
||||
# while `.felhom.yml` declares `required: false` and the API serves that verbatim. A
|
||||
# caller that trusts the contract gets a 400. Measured tonight on grafana; filed.
|
||||
if not f.get("required") and ty != "password":
|
||||
continue # the controller generates the optional secrets itself
|
||||
if ty == "path":
|
||||
p = f"{DRIVE}/{name}"
|
||||
values[ev] = p
|
||||
made.append(p)
|
||||
elif ty in ("secret", "password"):
|
||||
import secrets as _s
|
||||
values[ev] = "Drill-" + _s.token_hex(12)
|
||||
GENERATED.setdefault(name, {})[ev] = values[ev]
|
||||
elif f.get("default"):
|
||||
values[ev] = f["default"]
|
||||
else:
|
||||
values[ev] = f"drill-{name}"
|
||||
if made:
|
||||
guest("mkdir -p " + " ".join(made) + "; ls -ld " + " ".join(made))
|
||||
say(f" [1] made the drive paths this app requires: {made}")
|
||||
extra = [k for k in values if k not in ("DOMAIN", "SUBDOMAIN")]
|
||||
if extra:
|
||||
say(f" [1] required fields filled beyond DOMAIN/SUBDOMAIN: {extra}")
|
||||
return values
|
||||
|
||||
|
||||
def deploy(name, sub, extra_values=None):
|
||||
st = stack(name)
|
||||
if st.get("deployed"):
|
||||
say(f" [1] {name} already deployed — reusing")
|
||||
return True
|
||||
values = deploy_values(name, sub)
|
||||
if extra_values:
|
||||
values.update(extra_values)
|
||||
code, d = ctl("POST", f"/api/stacks/{name}/deploy", {"values": values})
|
||||
say(f" [1] deploy -> {code} {str(d)[:120]}")
|
||||
if code != "202":
|
||||
return False
|
||||
# WAIT FOR `deployed`, NOT FOR `running`. Measured 2026-09-21 on tandoor: docker reported the
|
||||
# container `healthy` while the controller's own state read `unhealthy` — a gate on `running`
|
||||
# alone therefore times out on an app that is up. The state is RECORDED rather than required;
|
||||
# the real gate is the fixture's own `wait_app`, which asks whether the APP answers.
|
||||
seen = None
|
||||
for _ in range(90):
|
||||
time.sleep(5)
|
||||
st = stack(name)
|
||||
seen = st.get("state")
|
||||
# `deployed` alone is NOT enough and `state` alone is NOT right. Measured 2026-09-21:
|
||||
# tandoor reads `unhealthy` while serving (R-618), so gating on "running" hangs; and romm
|
||||
# read `deployed=True, state=degraded, pinned_images=None` twenty seconds in, i.e. the
|
||||
# deploy had not finished writing app.yaml. The PIN is the deploy's own completion mark
|
||||
# (`runComposeDeploy` writes it), so that is what to wait for.
|
||||
pins = (st.get("app_config") or {}).get("pinned_images")
|
||||
if st.get("deployed") and pins and seen in ("running", "unhealthy", "degraded"):
|
||||
say(f" [1] deployed, controller state={seen}, "
|
||||
f"pinned={(st.get('app_config') or {}).get('pinned_images')}")
|
||||
if seen != "running":
|
||||
say(f" [1] NOTE: the controller's own state is {seen!r}, not 'running' — recorded, "
|
||||
f"not treated as a failure; the fixture's front-door wait is the real gate")
|
||||
return True
|
||||
say(f" [1] never became deployed (last controller state={seen!r})")
|
||||
return False
|
||||
|
||||
|
||||
def backup_now(name):
|
||||
code, d = ctl("POST", "/api/backup/run")
|
||||
say(f" [4] „Mentés most\" -> {code} {str(d)[:160]}")
|
||||
for _ in range(90):
|
||||
time.sleep(5)
|
||||
c2, s = ctl("GET", "/api/backup/status")
|
||||
dd = s.get("data") or {}
|
||||
if not dd.get("running", False):
|
||||
say(f" [4] backup idle; last={dd.get('last_run') or dd.get('last_db_dump')}")
|
||||
return True
|
||||
say(" [4] backup still running after 7.5 min — carrying on")
|
||||
return False
|
||||
|
||||
|
||||
def drill_bump(app, frm, to, service_hint=None):
|
||||
"""Commit the edge to the DRILL repo. catalog_since set by hand (the drill repo has no gates).
|
||||
|
||||
`frm`/`to` may be comma-separated lists of the SAME length: an app whose own version lives in
|
||||
two images (adventurelog's backend and frontend) moves both in one edge, while its engine
|
||||
sidecar stays where it is — `09` §3b Q3's rule is per SERVICE, and an app-half edge must move
|
||||
every service that carries the app's own version and no others.
|
||||
"""
|
||||
comp = f"{DRILL}/templates/{app}/docker-compose.yml"
|
||||
fy = f"{DRILL}/templates/{app}/.felhom.yml"
|
||||
s = open(comp).read()
|
||||
froms = [x.strip() for x in frm.split(",") if x.strip()]
|
||||
tos = [x.strip() for x in to.split(",") if x.strip()]
|
||||
if len(froms) != len(tos):
|
||||
say(f" [5] from/to lists differ in length: {froms} vs {tos}")
|
||||
return None
|
||||
for f1, t1 in zip(froms, tos):
|
||||
if f"image: {f1}" not in s:
|
||||
say(f" [5] FROM ref not found in compose: {f1}")
|
||||
return None
|
||||
s = s.replace(f"image: {f1}", f"image: {t1}")
|
||||
open(comp, "w").write(s)
|
||||
f = open(fy).read()
|
||||
today = datetime.now().strftime("%Y-%m-%d")
|
||||
f = re.sub(r'^catalog_since:.*$', f'catalog_since: "{today}"', f, count=1, flags=re.M)
|
||||
open(fy, "w").write(f)
|
||||
sh(["git", "-C", DRILL, "add", "-A"])
|
||||
sh(["git", "-C", DRILL, "commit", "-q", "-m", f"DRILL {app}: {frm} -> {to}"])
|
||||
r = sh(["git", "-C", DRILL, "push", "-q", "origin", "main"], timeout=120)
|
||||
h = sh(["git", "-C", DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip()
|
||||
say(f" [5] drill commit {h}: {app} {frm} -> {to} (push rc={r.returncode})")
|
||||
return h
|
||||
|
||||
|
||||
def sync_rescan(expect_app=None, expect_ref=None, tries=12, delay=5):
|
||||
"""Sync, rescan, and — when told what to expect — WAIT FOR THE BADGE TO CATCH UP.
|
||||
|
||||
R-607: `POST /api/sync` answers "nincs valtozas" while the catalog HAS moved, and
|
||||
`catalog_images` stays stale until a separate rescan. Tonight showed the rescan alone is not
|
||||
enough either: mealie's badge read "Naprakesz" seconds after its bump was pushed, and the
|
||||
Update that followed moved nothing and still reported "Frissitve". So when the caller knows
|
||||
which reference should appear, this polls for it and SAYS HOW LONG IT TOOK — which is the
|
||||
NUMBER R-607 asks for and has never had.
|
||||
"""
|
||||
t0 = time.time()
|
||||
ctl("POST", "/api/sync")
|
||||
time.sleep(2)
|
||||
ctl("POST", "/api/stacks/rescan")
|
||||
time.sleep(2)
|
||||
if not expect_app or not expect_ref:
|
||||
return None
|
||||
for i in range(tries):
|
||||
cat = stack(expect_app).get("catalog_images") or {}
|
||||
if expect_ref in cat.values():
|
||||
waited = round(time.time() - t0, 1)
|
||||
if i:
|
||||
say(f" [sync] the badge needed {waited}s and {i+1} sync+rescan rounds to catch up "
|
||||
f"to {expect_ref} — R-607's window, measured")
|
||||
return waited
|
||||
time.sleep(delay)
|
||||
ctl("POST", "/api/sync")
|
||||
time.sleep(1)
|
||||
ctl("POST", "/api/stacks/rescan")
|
||||
say(f" [sync] the badge NEVER caught up to {expect_ref} in {round(time.time()-t0,1)}s — "
|
||||
f"catalog_images = {stack(expect_app).get('catalog_images')}")
|
||||
return None
|
||||
|
||||
|
||||
def badges(name):
|
||||
out = {}
|
||||
for lang, suffix in (("hu", ""), ("en", "?lang=en")):
|
||||
h = page(f"/apps/{name}{suffix}")
|
||||
m = re.findall(r'<span class="tag tag-[^"]*"[^>]*title="([^"]*)"[^>]*>([^<]*)<', h)
|
||||
out[lang] = [{"title": a.strip(), "text": b.strip()} for a, b in m][:3]
|
||||
return out
|
||||
|
||||
|
||||
def press_update(name, poll=1.0, cap_s=1800):
|
||||
code, d = ctl("POST", f"/api/stacks/{name}/update")
|
||||
say(f" [6] Update -> {code} {str(d)[:220]}")
|
||||
if code not in ("202", "200"):
|
||||
return {"accepted": False, "http": code, "refusal": d, "phases": [], "duration_s": 0}
|
||||
phases, seen, t0 = [], None, time.time()
|
||||
while time.time() - t0 < cap_s:
|
||||
st = stack(name)
|
||||
ph = st.get("update_phase")
|
||||
if ph != seen:
|
||||
seen = ph
|
||||
rec = {"t": round(time.time() - t0, 1), "phase": ph,
|
||||
"label": st.get("update_phase_label"), "updating": st.get("updating"),
|
||||
"error": st.get("update_error"), "hold": st.get("hold_reason")}
|
||||
phases.append(rec)
|
||||
say(f" +{rec['t']:>6.1f}s phase={ph} label={rec['label']} "
|
||||
f"err={rec['error']} hold={rec['hold']}")
|
||||
if not st.get("updating") and ph in ("done", "failed", None) and time.time() - t0 > 3:
|
||||
break
|
||||
time.sleep(poll)
|
||||
st = stack(name)
|
||||
return {"accepted": True, "http": code, "phases": phases,
|
||||
"duration_s": round(time.time() - t0, 1),
|
||||
"final_phase": st.get("update_phase"), "update_error": st.get("update_error"),
|
||||
"hold_reason": st.get("hold_reason"), "state": st.get("state")}
|
||||
|
||||
|
||||
def observables(name):
|
||||
st = stack(name)
|
||||
ac = st.get("app_config") or {}
|
||||
live = guest(f"""
|
||||
grep -E '^\\s+image:' /opt/docker/stacks/{name}/docker-compose.yml 2>/dev/null | sed 's/^ *//'
|
||||
echo '---inspect---'
|
||||
for c in $(docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'); do
|
||||
echo -n "$c "; docker inspect "$c" --format '{{{{.Config.Image}}}} running={{{{.State.Running}}}} restarts={{{{.RestartCount}}}}'
|
||||
done
|
||||
""")
|
||||
a, _, b = live.partition("---inspect---")
|
||||
return {
|
||||
"pinned_images": ac.get("pinned_images"),
|
||||
"installed_images": {k: (v.get("ref") if isinstance(v, dict) else v)
|
||||
for k, v in (ac.get("installed_images") or {}).items()},
|
||||
"catalog_images": st.get("catalog_images"),
|
||||
"live_compose_image_lines": [x for x in a.strip().splitlines() if x.strip()],
|
||||
"docker_inspect": [x for x in b.strip().splitlines() if x.strip()],
|
||||
}
|
||||
|
||||
|
||||
def app_logs(name, lines=400):
|
||||
"""The app's own container log, DECODED. The endpoint answers a JSON envelope whose `logs` is
|
||||
one string with escaped newlines — a scan over the envelope sees a single enormous line and
|
||||
finds nothing, which reads exactly like "the app printed no migration line" and is not. R-96
|
||||
rule 3 in a new place: an absent line is not evidence when the instrument cannot see lines."""
|
||||
code, d = ctl("GET", f"/api/stacks/{name}/logs?lines={lines}")
|
||||
if isinstance(d, dict):
|
||||
data = d.get("data")
|
||||
if isinstance(data, dict) and isinstance(data.get("logs"), str):
|
||||
return data["logs"]
|
||||
if isinstance(d.get("_raw"), str):
|
||||
return d["_raw"]
|
||||
return str(d)
|
||||
|
||||
|
||||
def write_verdict(rec, appdir):
|
||||
os.makedirs(appdir, exist_ok=True)
|
||||
p = os.path.join(appdir, "verdict.json")
|
||||
json.dump(rec, open(p, "w"), indent=2, ensure_ascii=False)
|
||||
say(f" [9] verdict {rec['verdict']} -> {p}")
|
||||
|
||||
|
||||
def remove(name):
|
||||
"""Remove through the PRODUCT, never `docker rm` (live-probes rule). The remove endpoint
|
||||
refuses a running stack — `409 still running` — so the stop is part of the act, not a tidy-up."""
|
||||
c1, d1 = ctl("POST", f"/api/stacks/{name}/stop")
|
||||
say(f" [X] stop -> {c1} {str(d1)[:100]}")
|
||||
for _ in range(24):
|
||||
time.sleep(5)
|
||||
if stack(name).get("state") != "running":
|
||||
break
|
||||
code, d = ctl("POST", f"/api/stacks/{name}/remove",
|
||||
{"remove_hdd_data": True, "remove_backups": True})
|
||||
say(f" [X] remove (with drive data) -> {code} {str(d)[:160]}")
|
||||
if code == "409":
|
||||
# R-442's fail-closed guard: when the storage subsystem cannot RESOLVE the app's drive
|
||||
# path, the removal is REFUSED and the app is kept rather than half-deleted. On guest 9202
|
||||
# `/api/disks` answers `agent not configured`, so every app deployed with an HDD_PATH hits
|
||||
# this. The household's other choice — remove the app, KEEP the data — is accepted, and the
|
||||
# harness takes it, then tidies its own directory by name at teardown.
|
||||
say(" [X] refused because the drive path cannot be resolved (R-442, fail-closed and right)"
|
||||
" — removing the app and KEEPING the drive data instead")
|
||||
code, d = ctl("POST", f"/api/stacks/{name}/remove",
|
||||
{"remove_hdd_data": False, "remove_backups": True})
|
||||
say(f" [X] remove (keeping drive data) -> {code} {str(d)[:160]}")
|
||||
time.sleep(5)
|
||||
st = stack(name)
|
||||
left = guest(f"ls -d /opt/docker/stacks/{name} 2>/dev/null; "
|
||||
f"docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'")
|
||||
say(f" [X] after remove: deployed={st.get('deployed')} leftovers={left.strip()!r}")
|
||||
return code
|
||||
|
||||
|
||||
def app_env(name, key):
|
||||
"""Read one deploy value the CUSTOMER was given (e.g. the generated admin password) from the
|
||||
app's own `app.yaml`. This is not seeding — it is how the household logs in; the controller
|
||||
shows them the same value. Data still goes in through the app's own front door."""
|
||||
out = guest(f"grep -E '^\\s*{key}:' /opt/docker/stacks/{name}/app.yaml 2>/dev/null | head -1")
|
||||
if ":" in out:
|
||||
return out.split(":", 1)[1].strip().strip('"').strip("'")
|
||||
return ""
|
||||
|
||||
|
||||
def snapshots(name):
|
||||
"""The restorable copies the backups page offers for this app."""
|
||||
code, d = ctl("GET", f"/api/backup/snapshots?stack={name}")
|
||||
data = d.get("data") if isinstance(d, dict) else None
|
||||
if isinstance(data, dict):
|
||||
for k in ("snapshots", "items", "restore_points"):
|
||||
if isinstance(data.get(k), list):
|
||||
return data[k]
|
||||
return data if isinstance(data, list) else []
|
||||
|
||||
|
||||
def restore(name, snapshot_id=None, wait_s=1200):
|
||||
"""The household's own way out: the „Visszaállítás a mentésből" button on the backups page.
|
||||
|
||||
A FORM post, not an API call — `POST /backup/restore` with `_csrf`, `stack_name`,
|
||||
`snapshot_id` — because that is the button the sentence tells them to press.
|
||||
"""
|
||||
snaps = snapshots(name)
|
||||
if snapshot_id is None:
|
||||
if not snaps:
|
||||
say(f" [R] no restorable copy offered for {name}")
|
||||
return {"ok": False, "why": "no snapshot offered", "snapshots": snaps}
|
||||
first = snaps[0]
|
||||
snapshot_id = first.get("id") or first.get("snapshot_id") or first.get("short_id")
|
||||
say(f" [R] restoring {name} from snapshot {snapshot_id!r} (of {len(snaps)} offered)")
|
||||
sess = open(f"{SC}/sess.txt").read().strip()
|
||||
csrf = open(f"{SC}/csrf.txt").read().strip()
|
||||
r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", HOSTHDR, "-H", f"Cookie: {sess}",
|
||||
"-X", "POST",
|
||||
"--data-urlencode", f"_csrf={csrf}",
|
||||
"--data-urlencode", f"stack_name={name}",
|
||||
"--data-urlencode", f"snapshot_id={snapshot_id}",
|
||||
f"{BASE}/backup/restore"], timeout=180)
|
||||
head = (r.stdout or "").split("\n")[0].strip()
|
||||
loc = [l for l in (r.stdout or "").split("\n") if l.lower().startswith("location:")]
|
||||
say(f" [R] POST /backup/restore -> {head} {loc[:1]}")
|
||||
t0 = time.time()
|
||||
last = None
|
||||
while time.time() - t0 < wait_s:
|
||||
code, d = ctl("GET", "/api/backup/restore-status")
|
||||
dd = d.get("data") or {}
|
||||
cur = (dd.get("running"), dd.get("phase") or dd.get("state"), dd.get("message"))
|
||||
if cur != last:
|
||||
say(f" +{round(time.time()-t0,1):>6.1f}s restore {cur}")
|
||||
last = cur
|
||||
if not dd.get("running", False) and time.time() - t0 > 5:
|
||||
break
|
||||
time.sleep(2)
|
||||
st = stack(name)
|
||||
say(f" [R] after restore: state={st.get('state')} hold={st.get('hold_reason')!r} "
|
||||
f"phase={st.get('update_phase')}")
|
||||
return {"ok": True, "snapshot_id": snapshot_id, "snapshots": snaps,
|
||||
"http": head, "location": loc[:1], "seconds": round(time.time() - t0, 1),
|
||||
"state_after": st.get("state"), "hold_after": st.get("hold_reason"),
|
||||
"observables_after": observables(name)}
|
||||
@@ -814,6 +814,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-642** | **[P3-LOW] `POST /api/stacks/{name}/start` answers 200 *Stack … start completed* while the app is crash-looping.** MEASURED 2026-09-23 on 9202 twice: docmost 0.95.0 and romm 5.0.0 started on data their newer versions had migrated — both refused and restarted in a loop (`Restarting (1)`, front door 404) behind a 200. The same false-green class as R-443 (closed for the Update) and R-635, on the Start. It matters now because an undo (R-637) must never read the start's return as success. Evidence: `audits/update-rulings-2026-09-23/README.md`, `docmost-43`, `romm-43`. | **OPEN — P3; owner: CC** |
|
||||
| **R-643** | **[P2-MEDIUM] The ruled chain leaves the automatic update leg AT MOST 15 MINUTES a night.** FOUND 2026-09-23 while writing the build plan for `09` §3 decision 11 (*updates after the off-site copy, before the full-system backup*). The off-site leg starts at W+105m (`cmd/controller/main.go:961`) and the full-system backup's gate opens at W+2h (`quiesce/quiesce.go:656`, span to W+6h); the legs are clock-scheduled, not chained. One step takes ~1 min when it works and ~2–6 min when it fails and is undone. Options and the recommendation (the full-system backup waits for the leg inside its own window; the leg stops starting steps at W+5h) are in `09` §6.4. | **WAITING-ON-OPERATOR — `09` §6.4's one open point; owner: CC once answered** |
|
||||
| **R-644** | **[P3-LOW] `gokapi` on scratch guest 9202 is crash-looping — 329 restarts by 2026-09-23 07:51 UTC, *password does not appear to be a SHA-1 hash* — and the controller still lists it deployed.** OBSERVED at the start of the 2026-09-23 session, not caused by it. The twenty-eight walk's teardown (2026-09-22) removed a `gokapi` container left by R-633 by name; a `gokapi` is running again, recorded `deployed: true`. Not investigated (scope). Likely the R-633/R-634 shape — a restore-then-remove race leaving a record — and a scratch-box fact, not a customer one; filed so the next drill does not read it as its own. | **OPEN — P3; owner: CC; investigate before the next drill on 9202** |
|
||||
| **R-645** | **[P3-LOW] Lifting an update hold by the operator CLI lets the recovery unit be re-captured with the FAILED new definition within seconds — the copy the hold sentence names is overwritten.** MEASURED 2026-09-23 on 9202 during the undo bake-off: docmost was held at 07:59:16Z after a failed 0.95.0 → 0.96.0 update; `--clear-restore-hold docmost` + the controller restart it requires ran at ~07:59:23Z, and at **07:59:26Z** the controller logged *Recovery unit captured for docmost* — the unit's `compose/docker-compose.yml` now named `docmost/docmost:0.96.0`, the version that had just failed. The hold sentence had pointed the household at that unit („saját meghajtó, … 09:55"). The hold is what keeps the nightly legs off a held app (`isHeld`, v0.238.1); once it is lifted by hand, the checksum-gated refresh sees a changed definition and captures it. **Who it hits:** an operator who lifts a hold to inspect or repair, before restoring. With the undo (R-637) a failed update no longer holds unless the undo also fails, so the path is rarer — it does not go away. Candidate shapes, none chosen: the CLI refuses to lift an UPDATE hold (only a restore lifts it); or the lift also puts the pin back; or the capture skips an app whose pin is not what it is running. Evidence: `audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt` (the invalid run) and README §"Three things". | **OPEN — P3; owner: CC** |
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
Clearing a row means the check was DONE and its result recorded in that R-row —
|
||||
|
||||
Reference in New Issue
Block a user