R-649 closed: a failed install removes what it started (operator ruling); floor 0.266.0
gates / gates (push) Successful in 29s
gates / gates (push) Successful in 29s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -14,6 +14,13 @@
|
||||
> language, one screen, no identifiers in the prose. Same subjects, different readers; merging them
|
||||
> would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`**
|
||||
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
|
||||
## Ruling 2026-09-23 (night) — a failed install removes what it started (R-649, controller v0.266.0)
|
||||
|
||||
Operator: "On failed install, the controller should clean up the containers." `runComposeDeploy`'s failure
|
||||
branch runs `compose down` WITHOUT `-v` (named volumes kept — a reinstall after „keep my data" finds it)
|
||||
before recording not-deployed; a failed `down` is logged and Remove clears the rest. „Failed" now means
|
||||
nothing runs. Proven live on 9202 (`audits/r649-2026-09-23/`). Floor 0.266.0.
|
||||
|
||||
## 2026-09-23 (late evening) — clean-up: R-634's cause, held apps, the OOM storm (controller v0.265.0, hub v0.121.0); floor 0.265.0
|
||||
|
||||
**R-634 mechanism (reproduced on demand on 9202):** the in-memory `Deployed` flag is true from the moment a
|
||||
|
||||
@@ -1,3 +1,16 @@
|
||||
# ADDENDUM (same session) — R-649, the operator's ruling: a failed install removes what it started (controller v0.266.0)
|
||||
|
||||
Ruling: "On failed install, the controller should clean up the containers." Built in `964ae75`:
|
||||
`runComposeDeploy`'s failure branch runs `compose down` (no `-v` — volumes kept) before recording
|
||||
not-deployed. Red-proof: the `down` removed → `TestR649_AFailedInstallRemovesWhatItStarted` fails with
|
||||
`calls=["up -d"]`. Live on 9202 (`audits/r649-2026-09-23/`): outline with a never-healthy redis (drill
|
||||
template only) failed after 102.7 s → 0 containers left, 3 volumes kept, log line
|
||||
`the failed deploy's containers were removed (volumes kept) — R-649`. Floor **0.266.0** read back; both demo
|
||||
boxes on it in 20 s. R-649 closed; **open rows 331 → 330.** Teardown: outline removed, images by name, config
|
||||
identical to the saved copy, live catalog, drill reset. Nothing else touched.
|
||||
|
||||
---
|
||||
|
||||
# REPORT — clean-up evening: R-634's cause, held apps, the OOM storm (controller v0.265.0, hub v0.121.0)
|
||||
|
||||
2026-09-23 (late evening). Baselines verified live: controller `0a3026180ae8` (v0.264.0), agent
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-09-23 (late evening) — clean-up evening. The "runs but not installed" fault has its cause found and fixed. The fleet has the new version.**
|
||||
**Updated 2026-09-23 (night) — clean-up evening, plus your install ruling. The "runs but not installed" fault has its cause found and fixed. The fleet has the new version.**
|
||||
|
||||
**Decisions I took on my own: none.** One question for you is below.
|
||||
|
||||
@@ -18,8 +18,8 @@
|
||||
|
||||
**What was not clean.** While I wrote one test, it created an empty storage volume on your own machine. I saw it at once, checked it was new and unused, and removed it. Nothing else was touched. I filed a row so tests cannot do this again.
|
||||
|
||||
**What needs you — one question.** When an install fails for its own reasons, some of its containers can stay running while the box says "not installed". The household can remove them, so nothing is stuck. What should happen?
|
||||
- **Remove what the failed install started (my pick):** "failed" then always means nothing runs. The household presses Install again.
|
||||
- **Leave it as it is:** nothing changes. The page can say "not installed" over running containers until the household removes them.
|
||||
**Your ruling is built.** A failed install now removes the containers it started, and keeps the app's data. Proven on the scratch machine: an install that failed left no containers behind and kept its data. The floor is 0.266.0, and both demo machines updated themselves within 20 seconds. The list is now 330 rows.
|
||||
|
||||
**What needs you: nothing.**
|
||||
|
||||
**Nothing on Peti's machine or the off-site box was touched. On your own machine, only the hub was updated, plus the test mistake above. The demo-hp machine was not touched by hand. The scratch machine is back to its three standing apps and the real catalogue.**
|
||||
|
||||
@@ -0,0 +1,12 @@
|
||||
git:
|
||||
branch: main
|
||||
repo_url: https://gitea.dooplex.hu/admin/app-catalog-drill.git
|
||||
sync_interval: 15m
|
||||
token: <redacted>
|
||||
username: "admin"
|
||||
hub:
|
||||
update:
|
||||
health_timeout: 90s
|
||||
|
||||
18:08:42 catalog cache: 8777978 DRILL outline: redis healthcheck always fails (R-649 live proof)
|
||||
|
||||
@@ -0,0 +1,7 @@
|
||||
18:08:49 deploy outline -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
|
||||
18:08:49 + 0.0s outline: deployed=True deploying=True state=not_deployed err='None'
|
||||
18:08:52 + 3.0s outline: deployed=True deploying=True state=deploying err='None'
|
||||
18:09:40 + 51.4s outline: deployed=True deploying=True state=degraded err='None'
|
||||
18:10:35 + 105.9s outline: deployed=False deploying=False state=degraded err='exit code 1\nstderr: Image outlinewiki/outline:1.9.1 Pulling \n 72c03230f136 Already exists 0B\n 1acb9de2b08e Already exists 0B\n ad5e2960acd1 Already exists 0B\n 0e08d948bf76 Already exists 0B\n 464987f021'
|
||||
18:10:41 + 111.9s outline: deployed=False deploying=False state=not_deployed err='exit code 1\nstderr: Image outlinewiki/outline:1.9.1 Pulling \n 72c03230f136 Already exists 0B\n 1acb9de2b08e Already exists 0B\n ad5e2960acd1 Already exists 0B\n 0e08d948bf76 Already exists 0B\n 464987f021'
|
||||
18:11:41 final: {"deployed": false, "deploying": false, "state": "not_deployed", "deploy_error": "exit code 1\nstderr: Image outlinewiki/outline:1.9.1 Pulling \n 72c03230f136 Already exists 0B\n 1acb9de2b08e Already exists 0B\n ad5e2960acd1 Already exists 0B\n 0e08d948bf76 Already exists 0B\n 464987f021cb Pulling fs layer 0B\n f25fb26220a8 Pulling fs layer 0B\n 4d027270bfe6 Pulling fs layer 0B\n d256567eb8f6 Pulling fs layer 0B\n bc77c6bc12ba Pulling fs layer 0B\n 167140e5ddd9 Pulling fs layer 0B\n b6b3029da283 Pulling fs layer 0B\n 5d5585bcfc4f Pulling fs layer 0B\n 5726d50585da Pulling fs layer 0B\n 4a58d71
|
||||
@@ -0,0 +1,12 @@
|
||||
== containers of outline (expect none)
|
||||
== volumes of outline (expect kept)
|
||||
outline_outline_data
|
||||
outline_outline_postgres_data
|
||||
outline_outline_redis_data
|
||||
== app.yaml
|
||||
deployed: false
|
||||
desired_state: running
|
||||
== controller log
|
||||
2026/09/23 16:10:32 deploy.go:424: [ERROR] [stacks] Stack outline deploy failed after 102.7s: exit code 1
|
||||
2026/09/23 16:10:32 deploy.go:434: [INFO] [stacks] Stack outline: the failed deploy's containers were removed (volumes kept) — R-649
|
||||
|
||||
@@ -0,0 +1,24 @@
|
||||
18:11:55 === the household's Remove on the failed install
|
||||
18:11:55 [X] stop -> 200 {'ok': True, 'message': 'Stack outline stop completed'}
|
||||
18:12:27 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'outline', 'volumes_removed': ['outline_outline_data', 'outline_outline_postgres_data', 'outline_outline_redis_data'], 'hdd_pat
|
||||
18:12:34 [X] after remove: deployed=False leftovers='/opt/docker/stacks/outline'
|
||||
18:12:36 outline volumes left: 0
|
||||
18:12:41 rmi outlinewiki/outline:1.9.1
|
||||
KEPT postgres:16-alpine (in use by 1)
|
||||
KEPT redis:7-alpine (in use by 1)
|
||||
rmi-0.265.0
|
||||
|
||||
git:
|
||||
branch: main
|
||||
repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git
|
||||
sync_interval: 15m
|
||||
token: <redacted>
|
||||
username: ""
|
||||
hub:
|
||||
0
|
||||
|
||||
18:13:32 catalog: cfcfe52 upgrade-test: watch memory after the readback (harness v2, R-635/R-462)
|
||||
|
||||
18:13:34 controller.yaml vs saved: IDENTICAL
|
||||
|
||||
18:13:34 deployed: ['gokapi', 'paperless-ngx', 'privatebin']
|
||||
@@ -0,0 +1,16 @@
|
||||
=== FLOOR -> 0.266.0, MinAgent 0.131.0 declared (2026-09-23 18:13:43)
|
||||
POST /configuration/global-floor -> 303 Location: /configuration?flash=floor_set
|
||||
--- read back from the hub:
|
||||
"min_controller_version" value="0.266.0"
|
||||
"min_agent" value="0.131.0"
|
||||
"min_agent" value="0.131.0"
|
||||
--- watching the two demo boxes (up to 3 min)
|
||||
+10s demo-hp 9201: gitea.dooplex.hu/admin/felhom-controller:0.265.0 Up 37 minutes (healthy) | demo-felhom 9201: gitea.dooplex.hu/admin/felhom-controller:0.265.0 Up 37 minutes (healthy)
|
||||
+20s demo-hp 9201: gitea.dooplex.hu/admin/felhom-controller:0.266.0 Up 8 seconds (healthy) | demo-felhom 9201: gitea.dooplex.hu/admin/felhom-controller:0.266.0 Up 10 seconds (healthy)
|
||||
--- hub log:
|
||||
2026/09/23 18:13:46 [INFO] managed floor SERVED for demo-felhom: floor 0.266.0, agent requirement "0.131.0" from declared (golden 0.258.0)
|
||||
2026/09/23 18:13:47 [INFO] managed floor SERVED for demo-hp: floor 0.266.0, agent requirement "0.131.0" from declared (golden 0.258.0)
|
||||
--- hosts page:
|
||||
demo-felhom-8363b5
|
||||
demo-hp-bb76ea
|
||||
drill-r50-0a4f9a
|
||||
@@ -0,0 +1,20 @@
|
||||
# R-649 — a failed install removes what it started (controller v0.266.0), 2026-09-23
|
||||
|
||||
**Operator ruling 2026-09-23:** "On failed install, the controller should clean up the containers."
|
||||
**Method: endpoint-level** on scratch guest 9202 (deploy through `POST /api/stacks/<n>/deploy`, state from
|
||||
`GET /api/stacks/<n>`), drill catalog: `outline` with its redis healthcheck set to `false` in the DRILL
|
||||
template only, so `compose up -d` fails after postgres and redis have started.
|
||||
|
||||
| proof | result | evidence |
|
||||
|---|---|---|
|
||||
| failed install | `deploy failed after 102.7s: exit code 1` → **`the failed deploy's containers were removed (volumes kept) — R-649`**; 0 containers of project `outline`; volumes `outline_outline_data`, `_postgres_data`, `_redis_data` kept; `deployed: false` | `20-*`, `21-*` |
|
||||
| the household's Remove afterwards | 200, volumes removed with the data | `30-*` |
|
||||
| floor | 0.266.0 / MinAgent 0.131.0 read back; demo-hp and demo-felhom on 0.266.0 in 20 s | `40-*` |
|
||||
|
||||
Unit: `TestR649_AFailedInstallRemovesWhatItStarted` (stub compose; PATH = the stub only, R-650) and its
|
||||
control `TestR649_ASuccessfulInstallIsNotTakenDown`. Red-proof: the `down` removed → `calls=["up -d"]`.
|
||||
|
||||
**Teardown:** outline removed through the product; `outlinewiki/outline:1.9.1` and controller 0.265.0 removed
|
||||
by name (postgres/redis images kept — in use by standing apps); `controller.yaml` identical to
|
||||
`.pre-r649`; live catalog `cfcfe52`; drill repo reset to live `main`. Host: nothing. Hub: floor only.
|
||||
9201 not touched (reached 0.266.0 by the floor).
|
||||
@@ -0,0 +1,266 @@
|
||||
#!/usr/bin/env python3
|
||||
"""bakeoff.py — `09` §3 decision 19: the two ways of keeping the undo's last-second copy, measured on
|
||||
the same three apps (docmost / PostgreSQL, romm / MariaDB, vikunja / SQLite in a volume), guest 9202,
|
||||
drill catalog.
|
||||
|
||||
F — copy the folder: the app's NAMED volumes copied with its containers stopped; put back on failure.
|
||||
D — dump and load, as fixed by the morning spike: completion marker first; PostgreSQL drops and
|
||||
recreates the dump's schemas inside the load's transaction; MariaDB drops every table, then loads.
|
||||
|
||||
EVIDENCE, NOT PRODUCT. Product acts go through the endpoints the UI invokes (deploy, backup, sync,
|
||||
rescan, update, start, remove). The undo has no product path yet, so it is done by hand — the hold is
|
||||
lifted with the operator CLI + a controller restart (the only exit that exists), and the product's own
|
||||
Start supplies the app's secrets (they are never decrypted here).
|
||||
|
||||
Usage: python3 bakeoff.py <stage> <app> stages: prep fcopy break update undoF undoD cutoff state
|
||||
"""
|
||||
import json, os, sys, time
|
||||
sys.path.insert(0, ".")
|
||||
import walk as w
|
||||
from spike import FX, SUB, load, save, ts, break_edge, docmost_seed_b, docmost_verify_b, romm_seed_b, \
|
||||
romm_verify_b, vik_seed_b, vik_verify_b
|
||||
|
||||
EDGE = {"docmost": ("docmost/docmost:0.95.0", "docmost/docmost:0.96.0", 3000, 3999),
|
||||
"romm": ("rommapp/romm:5.0.0", "rommapp/romm:5.3.0", 8080, 8999),
|
||||
"vikunja": ("vikunja/vikunja:2.3.0", "vikunja/vikunja:2.6.0", 3456, 3999)}
|
||||
UNIT = {"docmost": "/mnt/sys_drive/felhom-data/backups/primary/docmost",
|
||||
"vikunja": "/mnt/sys_drive/felhom-data/backups/primary/vikunja",
|
||||
"romm": "/mnt/felhom-drives/scratch_hdd/userdata/romm/backups/primary/romm"}
|
||||
SUFFIX = ".pre-undo"
|
||||
|
||||
|
||||
def seed_b(app, sub, A):
|
||||
return {"docmost": docmost_seed_b, "romm": romm_seed_b, "vikunja": vik_seed_b}[app](sub, A)
|
||||
|
||||
|
||||
def verify_b(app, sub, A, B):
|
||||
r = {"docmost": docmost_verify_b, "romm": romm_verify_b, "vikunja": vik_verify_b}[app](sub, A, B)
|
||||
return all(r.values()) if isinstance(r, dict) else r
|
||||
|
||||
|
||||
def db_state(app):
|
||||
"""The table count and the migration ledger, asked of the engine (or the SQLite file, read-only)."""
|
||||
if app == "docmost":
|
||||
return w.guest("docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1 | tr '\\n' ' '; "
|
||||
"docker exec docmost-postgres psql -U docmost -d docmost -Atc \"select count(*)||' ledger, newest '||max(name) from kysely_migration\" 2>&1").strip()
|
||||
if app == "romm":
|
||||
q = lambda sql: f"docker exec romm-db sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" romm' 2>&1 | tr '\\n' ' '"
|
||||
base_tables = q("select count(*) from information_schema.tables where table_schema=database() and table_type=0x42415345205441424C45")
|
||||
alembic = q("select version_num from alembic_version")
|
||||
return w.guest(f'echo "tables=$({base_tables}) alembic=$({alembic})"').strip()
|
||||
return w.guest("""python3 - <<'PY'
|
||||
import sqlite3
|
||||
c=sqlite3.connect('file:/var/lib/docker/volumes/vikunja_vikunja_db/_data/vikunja.db?mode=ro',uri=True)
|
||||
t=c.execute("select count(*) from sqlite_master where type='table'").fetchone()[0]
|
||||
m=c.execute("select count(*), max(id) from migration").fetchone()
|
||||
print(f"tables={t} ledger={m[0]} newest={m[1]}")
|
||||
PY""").strip()
|
||||
|
||||
|
||||
def volumes(app):
|
||||
return [v for v in w.guest(f"docker volume ls -q --filter label=com.docker.compose.project={app}").split()
|
||||
if not v.endswith(SUFFIX)]
|
||||
|
||||
|
||||
def containers(app):
|
||||
return w.guest(f"docker ps -a --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'").split()
|
||||
|
||||
|
||||
def stage_prep(app):
|
||||
s = {"app": app, "sub": SUB[app]}
|
||||
w.say(f"=== prep {app}: deploy at the drill FROM pin")
|
||||
w.say("deployed:", w.deploy(app, s["sub"]))
|
||||
s["seedA"] = FX[app].seed(w, s["sub"], w.say)
|
||||
w.say("C1 A:", FX[app].verify(w, s["sub"], s["seedA"], w.say))
|
||||
w.backup_now(app)
|
||||
w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60)
|
||||
s["seedB"] = seed_b(app, s["sub"], s["seedA"])
|
||||
w.say("B reads back:", verify_b(app, s["sub"], s["seedA"], s["seedB"]))
|
||||
s["volumes"] = volumes(app)
|
||||
w.say("named volumes:", s["volumes"])
|
||||
save(app, s)
|
||||
|
||||
|
||||
def stage_fcopy(app):
|
||||
"""F at safety-dump time: stop, copy every named volume (cp -a into a sibling volume, and
|
||||
separately a tar, for the rate), start. Measures the EXTRA downtime and the disk."""
|
||||
s = load(app)
|
||||
s["state_before_update"] = db_state(app)
|
||||
w.say("db state before the update:", s["state_before_update"])
|
||||
vols = s["volumes"]
|
||||
w.say(w.guest(f"cp /opt/docker/stacks/{app}/docker-compose.yml /root/pre-{app}-compose.yml; grep -m1 'image: {EDGE[app][0]}' /root/pre-{app}-compose.yml"))
|
||||
script = f"""
|
||||
set -u
|
||||
cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}')
|
||||
t0=$(date +%s.%N)
|
||||
docker stop $cs >/dev/null
|
||||
t1=$(date +%s.%N)
|
||||
for v in {' '.join(vols)}; do
|
||||
docker volume rm -f "$v{SUFFIX}" >/dev/null 2>&1
|
||||
docker volume create "$v{SUFFIX}" >/dev/null
|
||||
a=$(date +%s.%N)
|
||||
docker run --rm -v "$v":/from:ro -v "$v{SUFFIX}":/to alpine:3.20 sh -c 'cp -a /from/. /to/ && sync' || echo "COPY FAILED $v"
|
||||
b=$(date +%s.%N)
|
||||
bytes=$(docker run --rm -v "$v":/from:ro alpine:3.20 du -sb /from | cut -f1)
|
||||
files=$(docker run --rm -v "$v":/from:ro alpine:3.20 sh -c 'find /from | wc -l')
|
||||
cbytes=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 du -sb /to | cut -f1)
|
||||
cfiles=$(docker run --rm -v "$v{SUFFIX}":/to:ro alpine:3.20 sh -c 'find /to | wc -l')
|
||||
mkdir -p /root/bk; c=$(date +%s.%N)
|
||||
docker run --rm -v "$v":/vol:ro -v /root/bk:/out alpine:3.20 tar cf /out/$v.tar -C /vol . ; d=$(date +%s.%N)
|
||||
tb=$(stat -c %s /root/bk/$v.tar); rm -f /root/bk/$v.tar
|
||||
python3 -c "print('VOL $v bytes=$bytes files=$files | copy bytes=$cbytes files=$cfiles | cp -a %.2fs | tar %.2fs (%s B)' % ($b-$a, $d-$c, '$tb'))"
|
||||
done
|
||||
t2=$(date +%s.%N)
|
||||
docker start $cs >/dev/null
|
||||
t3=$(date +%s.%N)
|
||||
python3 -c "print('stop %.2fs copy(all, cp -a + the tar measurement) %.2fs start %.2fs' % ($t1-$t0, $t2-$t1, $t3-$t2))"
|
||||
df -B1 --output=avail /var/lib/docker | tail -1 | awk '{{printf "free on the docker root: %.2f GiB\\n", $1/2^30}}'
|
||||
"""
|
||||
out = w.guest(script, timeout=1800)
|
||||
w.say(out)
|
||||
t0 = time.time()
|
||||
w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2)
|
||||
w.say(f"front door answering again {round(time.time()-t0,1)}s after the start")
|
||||
s["fcopy"] = out
|
||||
s["fcopy_at"] = ts()
|
||||
save(app, s)
|
||||
|
||||
|
||||
def stage_break(app):
|
||||
s = load(app)
|
||||
frm, to, p0, p1 = EDGE[app]
|
||||
s["break_commit"] = break_edge(app, frm, to, p0, p1)
|
||||
w.say("badge caught up after", w.sync_rescan(app, to), "s")
|
||||
save(app, s)
|
||||
|
||||
|
||||
def stage_update(app):
|
||||
s = load(app)
|
||||
r = w.press_update(app, poll=1.0)
|
||||
w.say(json.dumps({k: r[k] for k in ("final_phase", "hold_reason", "state", "duration_s")}, ensure_ascii=False))
|
||||
s.setdefault("updates", []).append(r)
|
||||
save(app, s)
|
||||
|
||||
|
||||
LIFT = """docker exec felhom-controller /usr/local/bin/felhom-controller --clear-restore-hold {app} 2>&1 | grep -E 'CLEARED|no restore hold'
|
||||
docker restart felhom-controller >/dev/null; sleep 15"""
|
||||
|
||||
# The OLD definition comes from the bake-off's OWN pre-update copy (/root/pre-<app>-compose.yml, taken
|
||||
# in fcopy), NEVER from the recovery unit: measured 2026-09-23 10:59, the unit was re-captured 10 s
|
||||
# after the hold was lifted — with the NEW definition — and a pin-back that read it started the new
|
||||
# version again. That is R-639 seen live, and why the product undo keeps its own copies.
|
||||
PINBACK = """S=/opt/docker/stacks/{app}; P=/root/pre-{app}-compose.yml
|
||||
grep -m1 'image: .*{frm}' $P >/dev/null || {{ echo "PRE-UPDATE COPY MISSING OR WRONG: $P"; exit 1; }}
|
||||
cp $P $S/applied-compose.yml; cp $P $S/docker-compose.yml
|
||||
sed -i 's#^\\(\\s*{svc}: \\){to}$#\\1{frm}#' $S/app.yaml; sed -n '/^pinned_images:/,$p' $S/app.yaml | head -4"""
|
||||
|
||||
|
||||
def lift_and_pinback(app):
|
||||
frm, to, _, _ = EDGE[app]
|
||||
t = time.time()
|
||||
w.say(w.guest(LIFT.format(app=app)))
|
||||
w.say(w.guest(PINBACK.format(unit=UNIT[app], app=app, svc=app, frm=frm, to=to)))
|
||||
w.login()
|
||||
return round(time.time() - t, 1)
|
||||
|
||||
|
||||
def stage_undoF(app):
|
||||
s = load(app)
|
||||
w.say(f"=== undo by F: {app}")
|
||||
w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)")
|
||||
vols = s["volumes"]
|
||||
script = f"""
|
||||
t0=$(date +%s.%N)
|
||||
cs=$(docker ps -a --filter label=com.docker.compose.project={app} -q); [ -n "$cs" ] && docker stop $cs >/dev/null
|
||||
for v in {' '.join(vols)}; do
|
||||
docker run --rm -v "$v{SUFFIX}":/from:ro -v "$v":/to alpine:3.20 sh -c 'rm -rf /to/..?* /to/.[!.]* /to/* ; cp -a /from/. /to/ && sync' || echo "RESTORE FAILED $v"
|
||||
done
|
||||
python3 -c "import time;print('volumes put back in %.2fs' % (time.time()-$t0))"
|
||||
"""
|
||||
w.say(w.guest(script, timeout=1800))
|
||||
t0 = time.time()
|
||||
c, d = w.ctl("POST", f"/api/stacks/{app}/start")
|
||||
w.say("product start ->", c)
|
||||
up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}[app], want=("200",), tries=60, delay=2)
|
||||
th = round(time.time() - t0, 1)
|
||||
A = FX[app].verify(w, s["sub"], s["seedA"], w.say)
|
||||
B = verify_b(app, s["sub"], s["seedA"], s["seedB"])
|
||||
st = db_state(app)
|
||||
w.say(f"F RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})")
|
||||
s["undoF"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st}
|
||||
save(app, s)
|
||||
|
||||
|
||||
def stage_undoD(app):
|
||||
"""D, on the SECOND failed update: its safety dump (written by the product, phase 3) is the copy."""
|
||||
s = load(app)
|
||||
w.say(f"=== undo by D: {app}")
|
||||
w.say("hold lift + pin back took", lift_and_pinback(app), "s (not part of a product undo)")
|
||||
if app == "vikunja":
|
||||
w.say("vikunja has no database server: the product wrote NO safety dump (R-641). D's copy for it "
|
||||
"is a volume tar at safety-dump time, which is F with extra steps — recorded, not re-run.")
|
||||
return
|
||||
c, d = w.ctl("POST", f"/api/stacks/{app}/start") # the product supplies the env; the app half will refuse
|
||||
time.sleep(8)
|
||||
if app == "docmost":
|
||||
load_cmd = """{ echo 'DROP SCHEMA public CASCADE; CREATE SCHEMA public;'; cat "$D"; } | docker exec -i docmost-postgres psql -v ON_ERROR_STOP=1 --single-transaction -U docmost -d docmost > /root/dload.out 2>&1"""
|
||||
marker = "-- PostgreSQL database dump complete"
|
||||
else:
|
||||
load_cmd = """{ echo 'SET FOREIGN_KEY_CHECKS=0;'; docker exec romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" -N -e "select concat(\\"DROP TABLE IF EXISTS \\`\\",table_name,\\"\\`;\\") from information_schema.tables where table_schema=\\"romm\\" and table_type=\\"BASE TABLE\\""' ; cat "$D"; } | docker exec -i romm-db sh -c 'mariadb -uroot -p"$MYSQL_ROOT_PASSWORD" romm' > /root/dload.out 2>&1"""
|
||||
marker = "-- Dump completed"
|
||||
script = f"""
|
||||
D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql | head -1); echo "undo copy: $(basename $D) $(stat -c %s $D) B"
|
||||
docker stop {app} >/dev/null 2>&1; docker update --restart=no {app} >/dev/null
|
||||
t0=$(date +%s.%N)
|
||||
if tail -n 5 "$D" | grep -q -- '{marker}'; then echo "marker present"; else echo "MARKER ABSENT - refusing"; exit 0; fi
|
||||
{load_cmd}; echo "load rc=$?"; tail -2 /root/dload.out
|
||||
python3 -c "import time;print('marker check + load %.2fs' % (time.time()-$t0))"
|
||||
docker update --restart=unless-stopped {app} >/dev/null; docker start {app} >/dev/null
|
||||
"""
|
||||
w.say(w.guest(script, timeout=900))
|
||||
t0 = time.time()
|
||||
up = w.wait_app(s["sub"], {"docmost": "/", "romm": "/api/heartbeat"}[app], want=("200",), tries=60, delay=2)
|
||||
th = round(time.time() - t0, 1)
|
||||
A = FX[app].verify(w, s["sub"], s["seedA"], w.say)
|
||||
B = verify_b(app, s["sub"], s["seedA"], s["seedB"])
|
||||
st = db_state(app)
|
||||
w.say(f"D RESULT: healthy={up} after {th}s A={A} B={B} db now: {st} (before the update: {s['state_before_update']})")
|
||||
s["undoD"] = {"healthy": up, "health_s": th, "A": A, "B": B, "db": st}
|
||||
save(app, s)
|
||||
|
||||
|
||||
def stage_cutoff(app):
|
||||
"""The cut-off copy, both methods, DETECTED before anything is loaded or swapped.
|
||||
F: a copy killed half-way (timeout) — compared with the source and checked for the finished-marker
|
||||
the build would write only after cp exits 0. D: the safety dump cut in half — the marker check."""
|
||||
s = load(app)
|
||||
big = max(s["volumes"], key=lambda v: int(w.guest(f"docker run --rm -v {v}:/v:ro alpine:3.20 du -sb /v | cut -f1").strip() or 0))
|
||||
script = f"""
|
||||
v={big}
|
||||
docker volume rm -f $v.cut >/dev/null 2>&1; docker volume create $v.cut >/dev/null
|
||||
cs=$(docker ps --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'); docker stop $cs >/dev/null
|
||||
# Kill the copy CONTAINER, not the client (killing `docker run` leaves the container copying — measured
|
||||
# on docmost 2026-09-23, rc 137 with a complete copy and a marker).
|
||||
docker run -d --name cutcopy -v $v:/from:ro -v $v.cut:/to alpine:3.20 sh -c 'cp -a /from/. /to/ && touch /to/.felhom-copy-complete' >/dev/null
|
||||
sleep 0.02; docker kill cutcopy >/dev/null 2>&1; echo "copy container exit=$(docker wait cutcopy) (137 = killed)"; docker rm -f cutcopy >/dev/null 2>&1
|
||||
echo "source bytes=$(docker run --rm -v $v:/f:ro alpine:3.20 du -sb /f | cut -f1) cut copy bytes=$(docker run --rm -v $v.cut:/f:ro alpine:3.20 du -sb /f | cut -f1)"
|
||||
echo "finished-marker in the cut copy: $(docker run --rm -v $v.cut:/f:ro alpine:3.20 sh -c 'ls /f/.felhom-copy-complete 2>/dev/null | wc -l') -> F refuses to swap"
|
||||
docker volume rm -f $v.cut >/dev/null; docker start $cs >/dev/null
|
||||
D=$(ls -t {UNIT[app]}/db-dumps/pre-restore-*.sql 2>/dev/null | head -1)
|
||||
if [ -n "$D" ]; then head -c $(( $(stat -c %s $D) / 2 )) $D > /root/cut.sql
|
||||
echo "D: whole copy marker: $(tail -n5 $D | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') cut copy marker: $(tail -n5 /root/cut.sql | grep -cE -- '-- (PostgreSQL database dump complete|Dump completed)') -> D refuses to load"
|
||||
rm -f /root/cut.sql
|
||||
else echo "D: no safety dump exists for this app (R-641)"; fi
|
||||
"""
|
||||
w.say(f"=== cut-off copy: {app} (largest volume {big})")
|
||||
out = w.guest(script, timeout=600)
|
||||
w.say(out)
|
||||
s["cutoff"] = out
|
||||
save(app, s)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
w.login()
|
||||
{"prep": stage_prep, "fcopy": stage_fcopy, "break": stage_break, "update": stage_update,
|
||||
"undoF": stage_undoF, "undoD": stage_undoD, "cutoff": stage_cutoff,
|
||||
"state": lambda a: w.say(db_state(a))}[sys.argv[1]](sys.argv[2])
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,194 @@
|
||||
#!/usr/bin/env python3
|
||||
"""live.py — controller v0.263.0's undo, proven on guest 9202 through the endpoints the UI invokes.
|
||||
|
||||
Nothing here performs an undo: the PRODUCT does. This presses Update, reads GET /api/stacks/<n>,
|
||||
fetches the app page in both languages, and reads the seeds back through each app's own front door.
|
||||
Seed C is written immediately before each Update, after every backup — so only the undo's own
|
||||
last-second copy can bring it back.
|
||||
"""
|
||||
import json, re, sys, time, html as H
|
||||
sys.path.insert(0, ".")
|
||||
import walk as w
|
||||
from spike import load, save, ts
|
||||
import bakeoff as b
|
||||
|
||||
HEALTH = {"docmost": "/", "romm": "/api/heartbeat", "vikunja": "/api/v1/info"}
|
||||
|
||||
|
||||
def seed_c(app, s):
|
||||
return b.seed_b(app, s["sub"], s["seedA"])
|
||||
|
||||
|
||||
def page_lines(app):
|
||||
out = {}
|
||||
for lang in ("hu", "en"):
|
||||
h = H.unescape(w.page(f"/apps/{app}?lang={lang}"))
|
||||
m = re.search(r'data-update-undone="true">([^<]*)<', h)
|
||||
hold = re.search(r'data-held="true">([^<]*)<', h)
|
||||
out[lang] = {"undone_line": m.group(1).strip() if m else None, "hold": hold.group(1).strip() if hold else None}
|
||||
return out
|
||||
|
||||
|
||||
def press(app, poll=0.5, on_phase=None):
|
||||
code, d = w.ctl("POST", f"/api/stacks/{app}/update")
|
||||
w.say(f" Update -> {code} {str(d)[:140]}")
|
||||
phases, seen, t0 = [], None, time.time()
|
||||
while time.time() - t0 < 1500:
|
||||
try:
|
||||
st = w.stack(app)
|
||||
except Exception as e:
|
||||
st = {}
|
||||
ph = st.get("update_phase")
|
||||
if ph != seen and ph is not None:
|
||||
seen = ph
|
||||
phases.append((round(time.time() - t0, 1), ph, st.get("update_phase_label")))
|
||||
w.say(f" +{phases[-1][0]:6.1f}s phase={ph} label={st.get('update_phase_label')}")
|
||||
if on_phase and on_phase(ph):
|
||||
return phases, "interrupted"
|
||||
if st and not st.get("updating") and ph in ("done", "failed", "undone") and time.time() - t0 > 2:
|
||||
break
|
||||
time.sleep(poll)
|
||||
st = w.stack(app)
|
||||
w.say(f" END phase={st.get('update_phase')} err={st.get('update_error')!r} hold={st.get('hold_reason')!r}")
|
||||
return phases, st
|
||||
|
||||
|
||||
def readback(app, s, with_c=True):
|
||||
sub = s["sub"]
|
||||
w.wait_app(sub, HEALTH[app], want=("200",), tries=60, delay=2)
|
||||
A = b.FX[app].verify(w, sub, s["seedA"], w.say)
|
||||
B = b.verify_b(app, sub, s["seedA"], s["seedB"])
|
||||
C = b.verify_b(app, sub, s["seedA"], s["seedC"]) if with_c and s.get("seedC") else None
|
||||
return {"A": A, "B": B, "C": C}
|
||||
|
||||
|
||||
def stage_undo(app):
|
||||
s = load(app)
|
||||
w.say(f"=== {app}: live undo by the product")
|
||||
s["seedC"] = seed_c(app, s); s["seedC_at"] = ts()
|
||||
w.say(" seed C written right before the Update:", bool(s["seedC"]))
|
||||
before = b.db_state(app); w.say(" db before:", before)
|
||||
phases, st = press(app)
|
||||
obs = w.observables(app)
|
||||
w.say(" observables:", json.dumps(obs))
|
||||
rb = readback(app, s)
|
||||
after = b.db_state(app)
|
||||
w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} db after: {after} (before: {before})")
|
||||
pl = page_lines(app); w.say(" PAGE:", json.dumps(pl, ensure_ascii=False))
|
||||
w.say(" leftover copies:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app} | wc -l").strip())
|
||||
s["live_undo"] = {"phases": phases, "end": {k: st.get(k) for k in ("update_phase", "update_error", "hold_reason")},
|
||||
"readback": rb, "db_before": before, "db_after": after, "page": pl, "obs": obs}
|
||||
save(app, s)
|
||||
|
||||
|
||||
|
||||
|
||||
def set_box_language(lang):
|
||||
"""POST /settings/language — the household's language switch (form + the dashboard's CSRF)."""
|
||||
sess = open(f"{w.SC}/sess.txt").read().strip(); csrf = open(f"{w.SC}/csrf.txt").read().strip()
|
||||
r = w.sh(["curl", "-sk", "-o", "/dev/null", "-w", "%{http_code}", "-H", w.HOSTHDR, "-H", f"Cookie: {sess}",
|
||||
"-H", f"X-CSRF-Token: {csrf}", "--data-urlencode", f"lang={lang}", "--data-urlencode", f"gorilla.csrf.Token={csrf}",
|
||||
f"{w.BASE}/settings/language"])
|
||||
w.say(f" box language -> {lang}: http {r.stdout.strip()}")
|
||||
|
||||
|
||||
def stage_powercut(app):
|
||||
"""Press Update; the moment the phase reads `undoing`, cut the guest's power (`pct stop`, a hard
|
||||
stop); boot it again and let the controller resume the undo."""
|
||||
import subprocess
|
||||
s = load(app)
|
||||
w.say(f"=== {app}: power cut DURING the undo")
|
||||
s["seedC"] = seed_c(app, s)
|
||||
cut = {}
|
||||
|
||||
def on_phase(ph):
|
||||
if ph == "undoing":
|
||||
t = time.time()
|
||||
r = subprocess.run(["ssh", "demo-hp", "pct stop 9202"], capture_output=True, text=True, timeout=120)
|
||||
cut["at"] = ts(); cut["rc"] = r.returncode
|
||||
w.say(f" >>> POWER CUT (pct stop 9202) in phase undoing: rc={r.returncode} in {round(time.time()-t,1)}s")
|
||||
return True
|
||||
return False
|
||||
phases, _ = press(app, poll=0.3, on_phase=on_phase)
|
||||
if not cut:
|
||||
w.say(" the cut never landed in `undoing` — recorded as a MISS"); return
|
||||
r = subprocess.run(["ssh", "demo-hp", "pct start 9202"], capture_output=True, text=True, timeout=180)
|
||||
w.say(f" guest started again: rc={r.returncode}")
|
||||
for i in range(60):
|
||||
time.sleep(5)
|
||||
try:
|
||||
w.login(); st = w.stack(app)
|
||||
if st:
|
||||
break
|
||||
except SystemExit:
|
||||
continue
|
||||
w.say(w.guest("docker logs felhom-controller 2>&1 | grep -E 'update recovery|resuming the UNDO|UNDONE|UNDO failed' | head -6"))
|
||||
t0 = time.time()
|
||||
while time.time() - t0 < 600:
|
||||
st = w.stack(app)
|
||||
if not st.get("updating") and st.get("update_phase") in ("undone", "failed"):
|
||||
break
|
||||
time.sleep(3)
|
||||
w.say(f" after the restart: phase={st.get('update_phase')} hold={st.get('hold_reason')!r}")
|
||||
rb = readback(app, s)
|
||||
w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} db: {b.db_state(app)}")
|
||||
w.say(" leftover copies:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app} | wc -l").strip())
|
||||
s["powercut"] = {"phases": phases, "cut": cut, "end": st.get("update_phase"), "readback": rb}
|
||||
save(app, s)
|
||||
|
||||
|
||||
def stage_cutoff(app):
|
||||
"""Press Update; while the NEW version is in `verifying`, take the finished-marker away from one of
|
||||
the undo copies (a copy cut off mid-way looks exactly like this). The undo must refuse to pour it
|
||||
back and HOLD with the new prefix."""
|
||||
s = load(app)
|
||||
w.say(f"=== {app}: a cut-off undo copy")
|
||||
done = {}
|
||||
|
||||
def on_phase(ph):
|
||||
if ph == "verifying" and not done:
|
||||
out = w.guest(f"""c=$(docker volume ls -q --filter label=felhom.undo-copy-of={app} | head -1); echo "copy: $c"
|
||||
docker run --rm -v $c:/c alpine sh -c 'ls -la /c; rm -f /c/felhom-undo-complete; ls /c'""")
|
||||
done["out"] = out
|
||||
w.say(" >>> finished-marker removed from one copy:\n" + out)
|
||||
return False
|
||||
phases, st = press(app, poll=0.5, on_phase=on_phase)
|
||||
before = b.db_state(app)
|
||||
pl = page_lines(app)
|
||||
w.say(" PAGE (box language hu):", json.dumps(pl, ensure_ascii=False))
|
||||
set_box_language("en")
|
||||
pl_en = page_lines(app)
|
||||
w.say(" PAGE (box language en):", json.dumps(pl_en, ensure_ascii=False))
|
||||
set_box_language("hu")
|
||||
w.say(" copies kept:", w.guest(f"docker volume ls -q --filter label=felhom.undo-copy-of={app}"))
|
||||
s["cutoff_live"] = {"phases": phases, "end": {k: st.get(k) for k in ("update_phase", "hold_reason")}, "page_hu_box": pl, "page_en_box": pl_en, "marker": done}
|
||||
save(app, s)
|
||||
|
||||
|
||||
def stage_fixprobe_and_press(app):
|
||||
"""After an undo: the catalog fixes the new version's probe; a PERSON presses Update; it must work
|
||||
and end the undone note."""
|
||||
s = load(app)
|
||||
frm, to, p0, p1 = b.EDGE[app]
|
||||
fy = f"{w.DRILL}/templates/{app}/.felhom.yml"
|
||||
f = open(fy).read().replace(f"port: {p1}", f"port: {p0}", 1); open(fy, "w").write(f)
|
||||
w.sh(["git", "-C", w.DRILL, "commit", "-qam", f"DRILL {app}: the probe fixed (port {p0}) — the step is now good"])
|
||||
w.sh(["git", "-C", w.DRILL, "push", "-q", "origin", "main"], timeout=120)
|
||||
w.sync_rescan(app, to)
|
||||
time.sleep(20)
|
||||
w.ctl("POST", "/api/sync"); time.sleep(3); w.ctl("POST", "/api/stacks/rescan")
|
||||
w.say(f"=== {app}: a person presses Update again after the undo (probe fixed in the catalog)")
|
||||
w.say(" before, the page:", json.dumps(page_lines(app), ensure_ascii=False))
|
||||
phases, st = press(app)
|
||||
rb = readback(app, s)
|
||||
w.say(f" READBACK A={rb['A']} B={rb['B']} C={rb['C']} installed={w.observables(app)['installed_images']}")
|
||||
pl = page_lines(app); w.say(" after, the page:", json.dumps(pl, ensure_ascii=False))
|
||||
w.say(" app.yaml last_update_undone:", w.guest(f"grep -c last_update_undone /opt/docker/stacks/{app}/app.yaml"))
|
||||
s["manual_after_undo"] = {"phases": phases, "end": st.get("update_phase"), "readback": rb, "page": pl}
|
||||
save(app, s)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
w.login()
|
||||
{"undo": stage_undo, "powercut": stage_powercut, "cutoff": stage_cutoff,
|
||||
"fixpress": stage_fixprobe_and_press}[sys.argv[1]](sys.argv[2])
|
||||
@@ -0,0 +1,60 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Point guest 9202 at the drill catalog (and a 90 s health timeout), or restore the saved config.
|
||||
|
||||
`09` §6.5: `git.repo_url` alone is INERT (R-615) — the cache dir must go too. The saved copy is
|
||||
`controller.yaml.pre-r649` (NOT the older `.pre-28`, which a restore must never pick up).
|
||||
"""
|
||||
import re, sys, io
|
||||
sys.path.insert(0, '.')
|
||||
import walk as w
|
||||
|
||||
VOL = "/var/lib/docker/volumes/felhom-controller-data/_data"
|
||||
DRILL_REPO = "https://gitea.dooplex.hu/admin/app-catalog-drill.git"
|
||||
|
||||
|
||||
def creds():
|
||||
for l in io.open("/home/kisfenyo/.git-credentials").read().strip().split("\n"):
|
||||
m = re.match(r'https://(admin):([^@]+)@gitea\.dooplex\.hu', l)
|
||||
if m:
|
||||
return m.group(1), m.group(2)
|
||||
raise SystemExit("no admin credential")
|
||||
|
||||
|
||||
def to_drill():
|
||||
u, t = creds()
|
||||
print(w.guest(f"""
|
||||
set -e
|
||||
test -f {VOL}/controller.yaml.pre-r649 || cp -p {VOL}/controller.yaml {VOL}/controller.yaml.pre-r649
|
||||
python3 - <<'PY'
|
||||
import re
|
||||
p = "{VOL}/controller.yaml"
|
||||
s = open(p).read()
|
||||
s = re.sub(r'(^\\s+repo_url: ).*$', r'\\g<1>{DRILL_REPO}', s, count=1, flags=re.M)
|
||||
s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+token: ).*$', r'\\g<1>"{t}"', s, count=1, flags=re.M)
|
||||
s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+username: ).*$', r'\\g<1>"{u}"', s, count=1, flags=re.M)
|
||||
if not re.search(r'^update:', s, re.M):
|
||||
s += "update:\\n health_timeout: 90s\\n"
|
||||
open(p, "w").write(s)
|
||||
PY
|
||||
rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache
|
||||
docker restart felhom-controller >/dev/null
|
||||
sleep 15
|
||||
grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: <redacted>/'
|
||||
grep -A2 '^update:' {VOL}/controller.yaml
|
||||
"""))
|
||||
|
||||
|
||||
def restore():
|
||||
print(w.guest(f"""
|
||||
set -e
|
||||
cp -p {VOL}/controller.yaml.pre-r649 {VOL}/controller.yaml
|
||||
rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache
|
||||
docker restart felhom-controller >/dev/null
|
||||
sleep 15
|
||||
grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: <redacted>/'
|
||||
grep -c '^update:' {VOL}/controller.yaml || true
|
||||
"""))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
to_drill() if sys.argv[1] == "drill" else restore()
|
||||
@@ -0,0 +1,160 @@
|
||||
#!/usr/bin/env python3
|
||||
"""spike.py — Part 1 of the 2026-09-23 brief: the AUTOMATIC UNDO, performed BY HAND on guest 9202.
|
||||
|
||||
EVIDENCE, NOT PRODUCT. Every product act goes through the endpoints the UI invokes (deploy, backup,
|
||||
sync, rescan, update, remove). The UNDO itself has no product path yet (that is what is being
|
||||
spiked), so it is performed by hand with plain docker/compose inside the guest, using exactly the
|
||||
steps the product would take, each one timed.
|
||||
|
||||
State between stages lives in state-<app>.json so each stage can be run, read, and only then
|
||||
followed by the next (the hand undo needs a person looking at what the previous step left).
|
||||
"""
|
||||
import json, os, sys, time, re
|
||||
sys.path.insert(0, ".")
|
||||
import walk as w
|
||||
import fixtures as fx
|
||||
|
||||
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
FX = {"docmost": fx.Docmost(), "vikunja": fx.Vikunja(), "romm": fx.Romm()}
|
||||
SUB = {"docmost": "docs", "vikunja": "tasks", "romm": "arcade"}
|
||||
|
||||
|
||||
def st_path(app):
|
||||
return os.path.join(HERE, f"state-{os.environ.get('GUEST', '9202')}-{app}.json")
|
||||
|
||||
|
||||
def load(app):
|
||||
return json.load(open(st_path(app))) if os.path.exists(st_path(app)) else {}
|
||||
|
||||
|
||||
def save(app, s):
|
||||
json.dump(s, open(st_path(app), "w"), indent=2, ensure_ascii=False)
|
||||
|
||||
|
||||
def ts():
|
||||
return time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
|
||||
|
||||
|
||||
# ---- a SECOND seed, written AFTER the backup and BEFORE the update. It is the discriminator: only
|
||||
# the pre-pin safety dump can hold it — the backup tier copy was taken before it existed. So if it
|
||||
# reads back after the undo, the undo used the safety dump; if only A reads back, it used the tier.
|
||||
def docmost_seed_b(sub, A):
|
||||
jar = "/tmp/dm.jar"
|
||||
w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json",
|
||||
data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST")
|
||||
name = "drillB" + os.urandom(3).hex()
|
||||
rc, code, out = w.app_curl(sub, "/api/spaces/create", "-b", jar, "-H", "Content-Type: application/json",
|
||||
data=json.dumps({"name": name, "slug": name.lower()}), method="POST")
|
||||
w.say(f" docmost seed B: /api/spaces/create http={code} {out[:160]}")
|
||||
return {"space": name} if code in ("200", "201") else None
|
||||
|
||||
|
||||
def docmost_verify_b(sub, A, B):
|
||||
jar = "/tmp/dm.jar"
|
||||
rc, code, out = w.app_curl(sub, "/api/auth/login", "-c", jar, "-H", "Content-Type: application/json",
|
||||
data=json.dumps({"email": A["email"], "password": A["pw"]}), method="POST")
|
||||
if code not in ("200", "201"):
|
||||
w.say(f" docmost B: cannot log in (http={code})"); return False
|
||||
rc, code, out = w.app_curl(sub, "/api/spaces", "-b", jar, "-H", "Content-Type: application/json",
|
||||
data="{}", method="POST")
|
||||
ok = code in ("200", "201") and B["space"] in out
|
||||
neg = "drillBnever" in out
|
||||
w.say(f" docmost B: /api/spaces http={code} seeded-space-listed={ok} (negative control listed={neg})")
|
||||
return ok and not neg
|
||||
|
||||
|
||||
def break_edge(app, frm, to, port_from, port_to):
|
||||
"""The failing edge: a REAL migrating image move, plus — in the DRILL template only — the
|
||||
health probe pointed at a port the app does not answer. Both in one drill commit."""
|
||||
fy = f"{w.DRILL}/templates/{app}/.felhom.yml"
|
||||
f = open(fy).read()
|
||||
m = re.search(r"(healthcheck:\n(?:.*\n){0,8}?\s+port: )" + str(port_from) + r"\b", f)
|
||||
assert m, "probe port not found"
|
||||
f = f[:m.end() - len(str(port_from))] + str(port_to) + f[m.end():]
|
||||
open(fy, "w").write(f)
|
||||
h = w.drill_bump(app, frm, to)
|
||||
return h
|
||||
|
||||
|
||||
def pg_state(container, db, user):
|
||||
return w.guest(f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from information_schema.tables where table_schema='public'\" 2>&1; "
|
||||
f"docker exec {container} psql -U {user} -d {db} -Atc \"select name from kysely_migration order by name desc limit 3\" 2>&1; "
|
||||
f"docker exec {container} psql -U {user} -d {db} -Atc \"select count(*) from kysely_migration\" 2>&1")
|
||||
|
||||
|
||||
def romm_seed_b(sub, A):
|
||||
"""A SECOND RomM user, created by the first (admin) one — written after the backup."""
|
||||
jar, tok = FX["romm"]._csrf(w, sub)
|
||||
user = "drillb" + os.urandom(3).hex()
|
||||
rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}",
|
||||
"-u", f"{A['user']}:{A['pw']}", "-H", "Content-Type: application/json",
|
||||
data=json.dumps({"username": user, "email": f"{user}@gate.invalid",
|
||||
"password": "Drill-" + os.urandom(8).hex(), "role": "viewer"}),
|
||||
method="POST")
|
||||
w.say(f" romm seed B: POST /api/users (as the admin) http={code} {body[:120]}")
|
||||
return {"user": user} if code in ("200", "201") else None
|
||||
|
||||
|
||||
def romm_verify_b(sub, A, B):
|
||||
jar, tok = FX["romm"]._csrf(w, sub)
|
||||
rc, code, body = w.app_curl(sub, "/api/users", "-b", jar, "-H", f"x-csrftoken: {tok}",
|
||||
"-u", f"{A['user']}:{A['pw']}")
|
||||
ok = code == "200" and B["user"] in body
|
||||
neg = "drillbnever" in body
|
||||
w.say(f" romm B: GET /api/users http={code} seeded-user-listed={ok} (negative control listed={neg})")
|
||||
return ok and not neg
|
||||
|
||||
|
||||
def tree(paths):
|
||||
cmd = "; ".join(f"echo \"{p}: files=$(find {p} -type f 2>/dev/null | wc -l) sum=$(find {p} -type f -exec sha256sum {{}} + 2>/dev/null | sort | sha256sum | cut -c1-16)\"" for p in paths)
|
||||
return w.guest(cmd)
|
||||
|
||||
|
||||
def my_state(container="romm-db", db="romm"):
|
||||
# The root password is used INSIDE the container from its own env — it never leaves it.
|
||||
q = lambda sql: f"docker exec {container} sh -c 'mariadb -uroot -p\"$MYSQL_ROOT_PASSWORD\" -N -e \"{sql}\" {db}' 2>&1 | tr '\\n' ' '"
|
||||
return w.guest(f"""echo -n "tables=$({q("select count(*) from information_schema.tables where table_schema=database()")}) "
|
||||
echo -n "alembic=$({q("select version_num from alembic_version")}) "
|
||||
echo -n "users=$({q("select count(*) from users")})"
|
||||
""")
|
||||
|
||||
|
||||
def vik_seed_b(sub, A):
|
||||
"""A second project, created AFTER the backup — plus a task with a real ATTACHMENT (a file the
|
||||
app writes into its files volume), uploaded through the app's own attachment API."""
|
||||
tok, why = FX["vikunja"]._token(w, sub, A)
|
||||
title = "drillB-" + os.urandom(4).hex()
|
||||
rc, code, body = w.app_curl(sub, "/api/v1/projects", "-H", f"Authorization: Bearer {tok}",
|
||||
"-H", "Content-Type: application/json", data=json.dumps({"title": title}), method="PUT")
|
||||
if code not in ("200", "201"):
|
||||
w.say(f" vikunja B: project refused {code} {body[:120]}"); return None
|
||||
pid = json.loads(body)["id"]
|
||||
rc, code, body = w.app_curl(sub, f"/api/v1/projects/{pid}/tasks", "-H", f"Authorization: Bearer {tok}",
|
||||
"-H", "Content-Type: application/json", data=json.dumps({"title": "task-" + title}), method="PUT")
|
||||
tid = json.loads(body)["id"] if code in ("200", "201") else None
|
||||
content = "drill attachment " + os.urandom(8).hex()
|
||||
fn = "/tmp/vik-att.txt"; open(fn, "w").write(content)
|
||||
rc, code2, body2 = w.app_curl(sub, f"/api/v1/tasks/{tid}/attachments", "-H", f"Authorization: Bearer {tok}",
|
||||
"-F", f"files=@{fn}", method="PUT")
|
||||
w.say(f" vikunja seed B: project http=200 task={tid} attachment upload http={code2} {body2[:120]}")
|
||||
return {"title": title, "pid": pid, "tid": tid, "att": content}
|
||||
|
||||
|
||||
def vik_verify_b(sub, A, B):
|
||||
tok, why = FX["vikunja"]._token(w, sub, A)
|
||||
if not tok:
|
||||
w.say(f" vikunja B: cannot log in {why}"); return {"project": False, "attachment": False}
|
||||
rc, code, body = w.app_curl(sub, f"/api/v1/projects/{B['pid']}", "-H", f"Authorization: Bearer {tok}")
|
||||
proj = code == "200" and B["title"] in body
|
||||
rc, code, body = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments", "-H", f"Authorization: Bearer {tok}")
|
||||
att_ok = False
|
||||
try:
|
||||
atts = json.loads(body)
|
||||
if atts:
|
||||
aid = atts[0]["id"]
|
||||
rc, c3, b3 = w.app_curl(sub, f"/api/v1/tasks/{B['tid']}/attachments/{aid}", "-H", f"Authorization: Bearer {tok}")
|
||||
att_ok = c3 == "200" and B["att"] in b3
|
||||
except Exception as e:
|
||||
w.say(f" vikunja B: attachments list unreadable http={code} {body[:120]}")
|
||||
w.say(f" vikunja B: project readback={proj} attachment content readback={att_ok}")
|
||||
return {"project": proj, "attachment": att_ok}
|
||||
@@ -0,0 +1,494 @@
|
||||
#!/usr/bin/env python3
|
||||
"""walk.py — ONE app's full update walk on guest 9202, through the product's own endpoints.
|
||||
|
||||
EVIDENCE, NOT PRODUCT. It presses exactly the buttons a person presses:
|
||||
POST /api/stacks/<n>/deploy · POST /api/sync · POST /api/stacks/rescan
|
||||
POST /api/stacks/<n>/update · POST /api/stacks/<n>/remove
|
||||
and reads GET /api/stacks/<n>. No controller code exists for it.
|
||||
|
||||
The walk, per `09` §6.4 and the update-night brief §4:
|
||||
1 deploy from the DRILL catalog at the LIVE pin
|
||||
2 seed through the app's OWN front door (R-156: never a volume, never SQL)
|
||||
3 read the seed back <- control C1; a fixture that cannot prove itself proves nothing
|
||||
4 „Mentés most"
|
||||
5 commit the real one-step bump to the DRILL repo, sync, rescan, read the badge in BOTH languages
|
||||
6 press the guarded Update, record every phase with timestamps
|
||||
7 read the seed back through the front door
|
||||
8 the four version observables side by side
|
||||
9 write the verdict record in `09`'s JSON shape
|
||||
|
||||
`inconclusive` is a first-class verdict and is NEVER collapsed into `failed`.
|
||||
"""
|
||||
import argparse, json, os, re, subprocess, sys, time
|
||||
from datetime import datetime, timezone
|
||||
|
||||
SC = "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/6e5a1a3b-6d8c-4ee1-bc3f-c555eb3f7578/scratchpad"
|
||||
EV = "/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/r649-2026-09-23"
|
||||
DRILL = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill"
|
||||
# GUEST=9201 selects demo-hp's hub-enabled guest (the mail proof); default 9202, the scratch guest.
|
||||
GUEST = os.environ.get("GUEST", "9202")
|
||||
BASE = {"9202": "https://192.168.0.114", "9201": "https://192.168.0.138"}[GUEST]
|
||||
DOMAIN = os.environ.get("DOMAIN", "enkisfelhom.hu")
|
||||
HOSTHDR = f"Host: felhom.{DOMAIN}"
|
||||
HP = "demo-hp"
|
||||
|
||||
LOG = []
|
||||
|
||||
|
||||
def say(*a):
|
||||
line = " ".join(str(x) for x in a)
|
||||
ts = datetime.now().strftime("%H:%M:%S")
|
||||
print(f"{ts} {line}", flush=True)
|
||||
LOG.append(f"{ts} {line}")
|
||||
|
||||
|
||||
def sh(args, timeout=300, inp=None):
|
||||
try:
|
||||
return subprocess.run(args, capture_output=True, text=True, timeout=timeout, input=inp)
|
||||
except (subprocess.TimeoutExpired, OSError) as e:
|
||||
return subprocess.CompletedProcess(args, 124, "", f"{e}")
|
||||
|
||||
|
||||
def guest(script, timeout=600):
|
||||
"""Run a bash script inside guest 9202. Piped as a file — never as an argument (quoting)."""
|
||||
r = sh(["ssh", "-o", "ConnectTimeout=20", "-o", "StrictHostKeyChecking=accept-new", HP,
|
||||
f"cat > /tmp/w{GUEST}.sh; pct push {GUEST} /tmp/w{GUEST}.sh /tmp/w.sh >/dev/null 2>&1; "
|
||||
f"pct exec {GUEST} -- bash /tmp/w.sh; rm -f /tmp/w{GUEST}.sh"],
|
||||
timeout=timeout, inp=script)
|
||||
return r.stdout or ""
|
||||
|
||||
|
||||
def login():
|
||||
pw = open(f"{SC}/.ctlpw").read().strip()
|
||||
sh(["curl", "-sk", "-D", f"{SC}/hdr.txt", "-o", "/dev/null", "-H", HOSTHDR,
|
||||
"-X", "POST", "--data-urlencode", f"password={pw}", f"{BASE}/login"])
|
||||
h = open(f"{SC}/hdr.txt").read()
|
||||
m = re.search(r"felhom_session=[A-Za-z0-9._-]+", h, re.I)
|
||||
if not m:
|
||||
sys.exit("login failed: no session cookie")
|
||||
open(f"{SC}/sess.txt", "w").write(m.group(0))
|
||||
r = sh(["curl", "-sk", "-L", "-H", HOSTHDR, "-H", f"Cookie: {m.group(0)}", f"{BASE}/"])
|
||||
c = re.search(r'<meta name="csrf-token" content="([^"]+)"', r.stdout or "")
|
||||
if not c:
|
||||
sys.exit("login failed: no csrf token")
|
||||
open(f"{SC}/csrf.txt", "w").write(c.group(1))
|
||||
|
||||
|
||||
def ctl(method, path, data=None, raw=False, tries=2):
|
||||
"""One controller API call. Re-logs in once on a 302/401 — the controller's session store is
|
||||
in memory, so any controller restart during the night invalidates it silently."""
|
||||
for attempt in range(tries):
|
||||
sess = open(f"{SC}/sess.txt").read().strip()
|
||||
csrf = open(f"{SC}/csrf.txt").read().strip()
|
||||
args = ["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", "-w", "\n%{http_code}"]
|
||||
if method != "GET":
|
||||
args += ["-H", f"X-CSRF-Token: {csrf}", "-H", "Content-Type: application/json",
|
||||
"-X", method]
|
||||
if data is not None:
|
||||
args += ["--data", json.dumps(data)]
|
||||
args.append(f"{BASE}{path}")
|
||||
r = sh(args)
|
||||
body, _, code = (r.stdout or "").rpartition("\n")
|
||||
if code.strip() in ("302", "401") and attempt + 1 < tries:
|
||||
login()
|
||||
continue
|
||||
if raw:
|
||||
return code.strip(), body
|
||||
try:
|
||||
return code.strip(), json.loads(body)
|
||||
except Exception:
|
||||
return code.strip(), {"_raw": body[:600]}
|
||||
return code.strip(), {"_raw": body[:600]}
|
||||
|
||||
|
||||
def page(path):
|
||||
sess = open(f"{SC}/sess.txt").read().strip()
|
||||
r = sh(["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", f"{BASE}{path}"])
|
||||
return r.stdout or ""
|
||||
|
||||
|
||||
def app_curl(sub, path, *extra, method=None, data=None, timeout=45):
|
||||
"""A call to the APP's own front door on 9202 — the household's route, not ours."""
|
||||
args = ["curl", "-sSk", "--max-time", str(timeout), "-H", f"Host: {sub}.{DOMAIN}",
|
||||
"-w", "\n%{http_code}"]
|
||||
if method:
|
||||
args += ["-X", method]
|
||||
if data is not None:
|
||||
args += ["--data-binary", "@-"]
|
||||
args += list(extra) + [f"{BASE}{path}"]
|
||||
r = sh(args, timeout=timeout + 30, inp=data)
|
||||
body, _, code = (r.stdout or "").rpartition("\n")
|
||||
return r.returncode, code.strip(), body
|
||||
|
||||
|
||||
def stack(name):
|
||||
_, d = ctl("GET", f"/api/stacks/{name}")
|
||||
return (d.get("data") or {}) if isinstance(d, dict) else {}
|
||||
|
||||
|
||||
def wait_app(sub, path="/", want=("200", "302", "303", "401", "403"), tries=60, delay=5):
|
||||
"""Settling says the container runs; this says the APP answers. Not the same thing."""
|
||||
last = None
|
||||
for _ in range(tries):
|
||||
rc, code, _ = app_curl(sub, path, timeout=15)
|
||||
last = (rc, code)
|
||||
if rc == 0 and code in want:
|
||||
return True
|
||||
time.sleep(delay)
|
||||
say(f" app never answered on {sub}{path} (last rc={last[0]} code={last[1]})")
|
||||
return False
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ the walk
|
||||
|
||||
|
||||
DRIVE = "/mnt/felhom-drives/scratch_hdd/userdata"
|
||||
|
||||
# What THIS run generated for a deploy, per app. Deploy secrets are ENCRYPTED AT REST in
|
||||
# `app.yaml` (`ENC:…`), which is right and which means a fixture cannot read an app's admin
|
||||
# password back off the box — the household sees it once. So the value the harness itself
|
||||
# generated is kept here for the life of the run, and nowhere else.
|
||||
GENERATED = {}
|
||||
|
||||
|
||||
def deploy_values(name, sub):
|
||||
"""Fill EVERY required deploy field the way the wizard would, by asking the box what this app
|
||||
asks for — `GET /api/stacks/<n>/deploy-fields` — instead of assuming DOMAIN+SUBDOMAIN.
|
||||
|
||||
Measured 2026-09-21: three apps in one batch refused at the deploy with a correct 400 because
|
||||
a required field was absent — `HDD_PATH` (navidrome, audiobookshelf) and an admin password
|
||||
(grafana). The refusals happen BEFORE anything is created (`deploy.go:324`), which is the only
|
||||
reason this was safe to discover by running it (live-probes rule).
|
||||
|
||||
A `path` field must name a directory that ALREADY EXISTS (`deploy.go:330`), so one is made on
|
||||
the scratch drive first — the same act the drive browser performs for a household.
|
||||
"""
|
||||
code, d = ctl("GET", f"/api/stacks/{name}/deploy-fields")
|
||||
fields = (((d.get("data") or {}).get("metadata") or {}).get("deploy_fields")) or []
|
||||
values = {"DOMAIN": DOMAIN, "SUBDOMAIN": sub}
|
||||
made = []
|
||||
for f in fields:
|
||||
ev, ty = f.get("env_var"), f.get("type")
|
||||
if ev in values:
|
||||
continue
|
||||
# `type: password` is MANDATORY whatever `required` says — `deploy.go:305-312` refuses
|
||||
# when the caller sends none, deliberately ("the user needs to know their password"),
|
||||
# while `.felhom.yml` declares `required: false` and the API serves that verbatim. A
|
||||
# caller that trusts the contract gets a 400. Measured tonight on grafana; filed.
|
||||
if not f.get("required") and ty != "password":
|
||||
continue # the controller generates the optional secrets itself
|
||||
if ty == "path":
|
||||
p = f"{DRIVE}/{name}"
|
||||
values[ev] = p
|
||||
made.append(p)
|
||||
elif ty in ("secret", "password"):
|
||||
import secrets as _s
|
||||
values[ev] = "Drill-" + _s.token_hex(12)
|
||||
GENERATED.setdefault(name, {})[ev] = values[ev]
|
||||
elif f.get("default"):
|
||||
values[ev] = f["default"]
|
||||
else:
|
||||
values[ev] = f"drill-{name}"
|
||||
if made:
|
||||
guest("mkdir -p " + " ".join(made) + "; ls -ld " + " ".join(made))
|
||||
say(f" [1] made the drive paths this app requires: {made}")
|
||||
extra = [k for k in values if k not in ("DOMAIN", "SUBDOMAIN")]
|
||||
if extra:
|
||||
say(f" [1] required fields filled beyond DOMAIN/SUBDOMAIN: {extra}")
|
||||
return values
|
||||
|
||||
|
||||
def deploy(name, sub, extra_values=None):
|
||||
st = stack(name)
|
||||
if st.get("deployed"):
|
||||
say(f" [1] {name} already deployed — reusing")
|
||||
return True
|
||||
values = deploy_values(name, sub)
|
||||
if extra_values:
|
||||
values.update(extra_values)
|
||||
code, d = ctl("POST", f"/api/stacks/{name}/deploy", {"values": values})
|
||||
say(f" [1] deploy -> {code} {str(d)[:120]}")
|
||||
if code != "202":
|
||||
return False
|
||||
# WAIT FOR `deployed`, NOT FOR `running`. Measured 2026-09-21 on tandoor: docker reported the
|
||||
# container `healthy` while the controller's own state read `unhealthy` — a gate on `running`
|
||||
# alone therefore times out on an app that is up. The state is RECORDED rather than required;
|
||||
# the real gate is the fixture's own `wait_app`, which asks whether the APP answers.
|
||||
seen = None
|
||||
for _ in range(90):
|
||||
time.sleep(5)
|
||||
st = stack(name)
|
||||
seen = st.get("state")
|
||||
# `deployed` alone is NOT enough and `state` alone is NOT right. Measured 2026-09-21:
|
||||
# tandoor reads `unhealthy` while serving (R-618), so gating on "running" hangs; and romm
|
||||
# read `deployed=True, state=degraded, pinned_images=None` twenty seconds in, i.e. the
|
||||
# deploy had not finished writing app.yaml. The PIN is the deploy's own completion mark
|
||||
# (`runComposeDeploy` writes it), so that is what to wait for.
|
||||
pins = (st.get("app_config") or {}).get("pinned_images")
|
||||
if st.get("deployed") and pins and seen in ("running", "unhealthy", "degraded"):
|
||||
say(f" [1] deployed, controller state={seen}, "
|
||||
f"pinned={(st.get('app_config') or {}).get('pinned_images')}")
|
||||
if seen != "running":
|
||||
say(f" [1] NOTE: the controller's own state is {seen!r}, not 'running' — recorded, "
|
||||
f"not treated as a failure; the fixture's front-door wait is the real gate")
|
||||
return True
|
||||
say(f" [1] never became deployed (last controller state={seen!r})")
|
||||
return False
|
||||
|
||||
|
||||
def backup_now(name):
|
||||
"""R-648 (2026-09-23): NO whole-box „Mentés most" from a drill, ever.
|
||||
|
||||
`POST /api/backup/run` is the only backup endpoint and it is WHOLE-BOX: on 9201 it stopped and
|
||||
restarted 9 of 10 standing apps twice, and on 9202 it broke a deploy in flight (R-634). The product
|
||||
has NO per-app backup endpoint (router.go: /backup/run, /backup/tier2 only); the per-app backup
|
||||
exists only inside the guarded update, whose `backing-up` phase calls RunAppBackupNow for the one
|
||||
app. So this presses nothing: the update takes the throwaway app's own backup, and says so in its
|
||||
phase list. A seed written "after the backup" is therefore written before the update's own backup
|
||||
— the undo's last-second copy is still the one that must bring it back."""
|
||||
say(f" [4] backup press SKIPPED for {name} (R-648: whole-box only; the update's backing-up phase backs up {name} alone)")
|
||||
return None
|
||||
|
||||
def drill_bump(app, frm, to, service_hint=None):
|
||||
"""Commit the edge to the DRILL repo. catalog_since set by hand (the drill repo has no gates).
|
||||
|
||||
`frm`/`to` may be comma-separated lists of the SAME length: an app whose own version lives in
|
||||
two images (adventurelog's backend and frontend) moves both in one edge, while its engine
|
||||
sidecar stays where it is — `09` §3b Q3's rule is per SERVICE, and an app-half edge must move
|
||||
every service that carries the app's own version and no others.
|
||||
"""
|
||||
comp = f"{DRILL}/templates/{app}/docker-compose.yml"
|
||||
fy = f"{DRILL}/templates/{app}/.felhom.yml"
|
||||
s = open(comp).read()
|
||||
froms = [x.strip() for x in frm.split(",") if x.strip()]
|
||||
tos = [x.strip() for x in to.split(",") if x.strip()]
|
||||
if len(froms) != len(tos):
|
||||
say(f" [5] from/to lists differ in length: {froms} vs {tos}")
|
||||
return None
|
||||
for f1, t1 in zip(froms, tos):
|
||||
if f"image: {f1}" not in s:
|
||||
say(f" [5] FROM ref not found in compose: {f1}")
|
||||
return None
|
||||
s = s.replace(f"image: {f1}", f"image: {t1}")
|
||||
open(comp, "w").write(s)
|
||||
f = open(fy).read()
|
||||
today = datetime.now().strftime("%Y-%m-%d")
|
||||
f = re.sub(r'^catalog_since:.*$', f'catalog_since: "{today}"', f, count=1, flags=re.M)
|
||||
open(fy, "w").write(f)
|
||||
sh(["git", "-C", DRILL, "add", "-A"])
|
||||
sh(["git", "-C", DRILL, "commit", "-q", "-m", f"DRILL {app}: {frm} -> {to}"])
|
||||
r = sh(["git", "-C", DRILL, "push", "-q", "origin", "main"], timeout=120)
|
||||
h = sh(["git", "-C", DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip()
|
||||
say(f" [5] drill commit {h}: {app} {frm} -> {to} (push rc={r.returncode})")
|
||||
return h
|
||||
|
||||
|
||||
def sync_rescan(expect_app=None, expect_ref=None, tries=12, delay=5):
|
||||
"""Sync, rescan, and — when told what to expect — WAIT FOR THE BADGE TO CATCH UP.
|
||||
|
||||
R-607: `POST /api/sync` answers "nincs valtozas" while the catalog HAS moved, and
|
||||
`catalog_images` stays stale until a separate rescan. Tonight showed the rescan alone is not
|
||||
enough either: mealie's badge read "Naprakesz" seconds after its bump was pushed, and the
|
||||
Update that followed moved nothing and still reported "Frissitve". So when the caller knows
|
||||
which reference should appear, this polls for it and SAYS HOW LONG IT TOOK — which is the
|
||||
NUMBER R-607 asks for and has never had.
|
||||
"""
|
||||
t0 = time.time()
|
||||
ctl("POST", "/api/sync")
|
||||
time.sleep(2)
|
||||
ctl("POST", "/api/stacks/rescan")
|
||||
time.sleep(2)
|
||||
if not expect_app or not expect_ref:
|
||||
return None
|
||||
for i in range(tries):
|
||||
cat = stack(expect_app).get("catalog_images") or {}
|
||||
if expect_ref in cat.values():
|
||||
waited = round(time.time() - t0, 1)
|
||||
if i:
|
||||
say(f" [sync] the badge needed {waited}s and {i+1} sync+rescan rounds to catch up "
|
||||
f"to {expect_ref} — R-607's window, measured")
|
||||
return waited
|
||||
time.sleep(delay)
|
||||
ctl("POST", "/api/sync")
|
||||
time.sleep(1)
|
||||
ctl("POST", "/api/stacks/rescan")
|
||||
say(f" [sync] the badge NEVER caught up to {expect_ref} in {round(time.time()-t0,1)}s — "
|
||||
f"catalog_images = {stack(expect_app).get('catalog_images')}")
|
||||
return None
|
||||
|
||||
|
||||
def badges(name):
|
||||
out = {}
|
||||
for lang, suffix in (("hu", ""), ("en", "?lang=en")):
|
||||
h = page(f"/apps/{name}{suffix}")
|
||||
m = re.findall(r'<span class="tag tag-[^"]*"[^>]*title="([^"]*)"[^>]*>([^<]*)<', h)
|
||||
out[lang] = [{"title": a.strip(), "text": b.strip()} for a, b in m][:3]
|
||||
return out
|
||||
|
||||
|
||||
def press_update(name, poll=1.0, cap_s=1800):
|
||||
code, d = ctl("POST", f"/api/stacks/{name}/update")
|
||||
say(f" [6] Update -> {code} {str(d)[:220]}")
|
||||
if code not in ("202", "200"):
|
||||
return {"accepted": False, "http": code, "refusal": d, "phases": [], "duration_s": 0}
|
||||
phases, seen, t0 = [], None, time.time()
|
||||
while time.time() - t0 < cap_s:
|
||||
st = stack(name)
|
||||
ph = st.get("update_phase")
|
||||
if ph != seen:
|
||||
seen = ph
|
||||
rec = {"t": round(time.time() - t0, 1), "phase": ph,
|
||||
"label": st.get("update_phase_label"), "updating": st.get("updating"),
|
||||
"error": st.get("update_error"), "hold": st.get("hold_reason")}
|
||||
phases.append(rec)
|
||||
say(f" +{rec['t']:>6.1f}s phase={ph} label={rec['label']} "
|
||||
f"err={rec['error']} hold={rec['hold']}")
|
||||
if not st.get("updating") and ph in ("done", "failed", None) and time.time() - t0 > 3:
|
||||
break
|
||||
time.sleep(poll)
|
||||
st = stack(name)
|
||||
return {"accepted": True, "http": code, "phases": phases,
|
||||
"duration_s": round(time.time() - t0, 1),
|
||||
"final_phase": st.get("update_phase"), "update_error": st.get("update_error"),
|
||||
"hold_reason": st.get("hold_reason"), "state": st.get("state")}
|
||||
|
||||
|
||||
def observables(name):
|
||||
st = stack(name)
|
||||
ac = st.get("app_config") or {}
|
||||
live = guest(f"""
|
||||
grep -E '^\\s+image:' /opt/docker/stacks/{name}/docker-compose.yml 2>/dev/null | sed 's/^ *//'
|
||||
echo '---inspect---'
|
||||
for c in $(docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'); do
|
||||
echo -n "$c "; docker inspect "$c" --format '{{{{.Config.Image}}}} running={{{{.State.Running}}}} restarts={{{{.RestartCount}}}}'
|
||||
done
|
||||
""")
|
||||
a, _, b = live.partition("---inspect---")
|
||||
return {
|
||||
"pinned_images": ac.get("pinned_images"),
|
||||
"installed_images": {k: (v.get("ref") if isinstance(v, dict) else v)
|
||||
for k, v in (ac.get("installed_images") or {}).items()},
|
||||
"catalog_images": st.get("catalog_images"),
|
||||
"live_compose_image_lines": [x for x in a.strip().splitlines() if x.strip()],
|
||||
"docker_inspect": [x for x in b.strip().splitlines() if x.strip()],
|
||||
}
|
||||
|
||||
|
||||
def app_logs(name, lines=400):
|
||||
"""The app's own container log, DECODED. The endpoint answers a JSON envelope whose `logs` is
|
||||
one string with escaped newlines — a scan over the envelope sees a single enormous line and
|
||||
finds nothing, which reads exactly like "the app printed no migration line" and is not. R-96
|
||||
rule 3 in a new place: an absent line is not evidence when the instrument cannot see lines."""
|
||||
code, d = ctl("GET", f"/api/stacks/{name}/logs?lines={lines}")
|
||||
if isinstance(d, dict):
|
||||
data = d.get("data")
|
||||
if isinstance(data, dict) and isinstance(data.get("logs"), str):
|
||||
return data["logs"]
|
||||
if isinstance(d.get("_raw"), str):
|
||||
return d["_raw"]
|
||||
return str(d)
|
||||
|
||||
|
||||
def write_verdict(rec, appdir):
|
||||
os.makedirs(appdir, exist_ok=True)
|
||||
p = os.path.join(appdir, "verdict.json")
|
||||
json.dump(rec, open(p, "w"), indent=2, ensure_ascii=False)
|
||||
say(f" [9] verdict {rec['verdict']} -> {p}")
|
||||
|
||||
|
||||
def remove(name):
|
||||
"""Remove through the PRODUCT, never `docker rm` (live-probes rule). The remove endpoint
|
||||
refuses a running stack — `409 still running` — so the stop is part of the act, not a tidy-up."""
|
||||
c1, d1 = ctl("POST", f"/api/stacks/{name}/stop")
|
||||
say(f" [X] stop -> {c1} {str(d1)[:100]}")
|
||||
for _ in range(24):
|
||||
time.sleep(5)
|
||||
if stack(name).get("state") != "running":
|
||||
break
|
||||
code, d = ctl("POST", f"/api/stacks/{name}/remove",
|
||||
{"remove_hdd_data": True, "remove_backups": True})
|
||||
say(f" [X] remove (with drive data) -> {code} {str(d)[:160]}")
|
||||
if code == "409":
|
||||
# R-442's fail-closed guard: when the storage subsystem cannot RESOLVE the app's drive
|
||||
# path, the removal is REFUSED and the app is kept rather than half-deleted. On guest 9202
|
||||
# `/api/disks` answers `agent not configured`, so every app deployed with an HDD_PATH hits
|
||||
# this. The household's other choice — remove the app, KEEP the data — is accepted, and the
|
||||
# harness takes it, then tidies its own directory by name at teardown.
|
||||
say(" [X] refused because the drive path cannot be resolved (R-442, fail-closed and right)"
|
||||
" — removing the app and KEEPING the drive data instead")
|
||||
code, d = ctl("POST", f"/api/stacks/{name}/remove",
|
||||
{"remove_hdd_data": False, "remove_backups": True})
|
||||
say(f" [X] remove (keeping drive data) -> {code} {str(d)[:160]}")
|
||||
time.sleep(5)
|
||||
st = stack(name)
|
||||
left = guest(f"ls -d /opt/docker/stacks/{name} 2>/dev/null; "
|
||||
f"docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'")
|
||||
say(f" [X] after remove: deployed={st.get('deployed')} leftovers={left.strip()!r}")
|
||||
return code
|
||||
|
||||
|
||||
def app_env(name, key):
|
||||
"""Read one deploy value the CUSTOMER was given (e.g. the generated admin password) from the
|
||||
app's own `app.yaml`. This is not seeding — it is how the household logs in; the controller
|
||||
shows them the same value. Data still goes in through the app's own front door."""
|
||||
out = guest(f"grep -E '^\\s*{key}:' /opt/docker/stacks/{name}/app.yaml 2>/dev/null | head -1")
|
||||
if ":" in out:
|
||||
return out.split(":", 1)[1].strip().strip('"').strip("'")
|
||||
return ""
|
||||
|
||||
|
||||
def snapshots(name):
|
||||
"""The restorable copies the backups page offers for this app."""
|
||||
code, d = ctl("GET", f"/api/backup/snapshots?stack={name}")
|
||||
data = d.get("data") if isinstance(d, dict) else None
|
||||
if isinstance(data, dict):
|
||||
for k in ("snapshots", "items", "restore_points"):
|
||||
if isinstance(data.get(k), list):
|
||||
return data[k]
|
||||
return data if isinstance(data, list) else []
|
||||
|
||||
|
||||
def restore(name, snapshot_id=None, wait_s=1200):
|
||||
"""The household's own way out: the „Visszaállítás a mentésből" button on the backups page.
|
||||
|
||||
A FORM post, not an API call — `POST /backup/restore` with `_csrf`, `stack_name`,
|
||||
`snapshot_id` — because that is the button the sentence tells them to press.
|
||||
"""
|
||||
snaps = snapshots(name)
|
||||
if snapshot_id is None:
|
||||
if not snaps:
|
||||
say(f" [R] no restorable copy offered for {name}")
|
||||
return {"ok": False, "why": "no snapshot offered", "snapshots": snaps}
|
||||
first = snaps[0]
|
||||
snapshot_id = first.get("id") or first.get("snapshot_id") or first.get("short_id")
|
||||
say(f" [R] restoring {name} from snapshot {snapshot_id!r} (of {len(snaps)} offered)")
|
||||
sess = open(f"{SC}/sess.txt").read().strip()
|
||||
csrf = open(f"{SC}/csrf.txt").read().strip()
|
||||
r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", HOSTHDR, "-H", f"Cookie: {sess}",
|
||||
"-X", "POST",
|
||||
"--data-urlencode", f"_csrf={csrf}",
|
||||
"--data-urlencode", f"stack_name={name}",
|
||||
"--data-urlencode", f"snapshot_id={snapshot_id}",
|
||||
f"{BASE}/backup/restore"], timeout=180)
|
||||
head = (r.stdout or "").split("\n")[0].strip()
|
||||
loc = [l for l in (r.stdout or "").split("\n") if l.lower().startswith("location:")]
|
||||
say(f" [R] POST /backup/restore -> {head} {loc[:1]}")
|
||||
t0 = time.time()
|
||||
last = None
|
||||
while time.time() - t0 < wait_s:
|
||||
code, d = ctl("GET", "/api/backup/restore-status")
|
||||
dd = d.get("data") or {}
|
||||
cur = (dd.get("running"), dd.get("phase") or dd.get("state"), dd.get("message"))
|
||||
if cur != last:
|
||||
say(f" +{round(time.time()-t0,1):>6.1f}s restore {cur}")
|
||||
last = cur
|
||||
if not dd.get("running", False) and time.time() - t0 > 5:
|
||||
break
|
||||
time.sleep(2)
|
||||
st = stack(name)
|
||||
say(f" [R] after restore: state={st.get('state')} hold={st.get('hold_reason')!r} "
|
||||
f"phase={st.get('update_phase')}")
|
||||
return {"ok": True, "snapshot_id": snapshot_id, "snapshots": snaps,
|
||||
"http": head, "location": loc[:1], "seconds": round(time.time() - t0, 1),
|
||||
"state_after": st.get("state"), "hold_after": st.get("hold_reason"),
|
||||
"observables_after": observables(name)}
|
||||
@@ -0,0 +1,26 @@
|
||||
"""Deploy one app (or none) and record every change of deployed/deploying/state with a timestamp."""
|
||||
import sys, time, json
|
||||
import walk as w
|
||||
app = sys.argv[1]; cap = int(sys.argv[2]) if len(sys.argv) > 2 else 900
|
||||
w.login()
|
||||
if "--deploy" in sys.argv:
|
||||
vals = w.deploy_values(app, app)
|
||||
c, d = w.ctl("POST", f"/api/stacks/{app}/deploy", {"values": vals} if vals else {})
|
||||
w.say(f"deploy {app} -> {c} {str(d)[:160]}")
|
||||
last = None; t0 = time.time()
|
||||
while time.time() - t0 < cap:
|
||||
st = w.stack(app)
|
||||
cur = (st.get("deployed"), st.get("deploying"), st.get("state"), st.get("deploy_error"))
|
||||
if cur != last:
|
||||
w.say(f"+{time.time()-t0:6.1f}s {app}: deployed={cur[0]} deploying={cur[1]} state={cur[2]} err={str(cur[3])[:200]!r}")
|
||||
last = cur
|
||||
if "--until-settled" in sys.argv and cur[1] is False and time.time() - t0 > 30 and cur[2] in ("running", "not_deployed", "error", "stopped", "exited"):
|
||||
# settled: keep watching 60 s more for late flips
|
||||
end = time.time() + 60
|
||||
while time.time() < end:
|
||||
st = w.stack(app); c2 = (st.get("deployed"), st.get("deploying"), st.get("state"), st.get("deploy_error"))
|
||||
if c2 != last: w.say(f"+{time.time()-t0:6.1f}s {app}: deployed={c2[0]} deploying={c2[1]} state={c2[2]} err={str(c2[3])[:200]!r}"); last = c2
|
||||
time.sleep(3)
|
||||
break
|
||||
time.sleep(3)
|
||||
w.say("final:", json.dumps({k: w.stack(app).get(k) for k in ("deployed","deploying","state","deploy_error","containers")}, default=str)[:600])
|
||||
@@ -357,4 +357,5 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
|
||||
| **R-636** | **Six hours of OOM kills sent the same single warning as one hiccup (P2).** v0.265.0 + hub v0.121.0: the kernel `oom_kill` counter (the `OOMKilled` flag is sticky and cannot count); ≥ 20 kills in 30 min of one container run → ONE `app_oom_storm` (error, operator-only, per-app cooldown). | **CLOSED 2026-09-23 — PROVEN-LIVE on the controller (`45-*`, `46-*`: RomM at 320M, storm at 21 kills, still one at 49); hub side by unit tests** | as above |
|
||||
| **R-647** | **Three leftovers of the update mail (P3).** v0.265.0: a held update's error is the key `update.error.held`, rendered per reader on both pages and the API; `copy_holds` travels as its key; the two log wordings fixed. | **CLOSED 2026-09-23 — PROVEN-LIVE for (1) (`43-*`, `44-*`); (2)(3) by red-proofed tests** | as above |
|
||||
| **R-648** | **The drill's „Mentés most" was whole-box (P3).** No per-app backup endpoint exists; the harness (`audits/cleanup-2026-09-23/walk.py` `backup_now`) now presses nothing and the guarded update's own `backing-up` phase backs up the throwaway app alone. | **CLOSED 2026-09-23 — PROVEN-LIVE (`43-*`: phase `backing-up` for vikunja only)** | as above |
|
||||
| **R-649** | **A failed install could leave containers under „not deployed" (P2).** Operator ruling 2026-09-23 (option a): the controller cleans up. Closed in **v0.266.0** (`964ae75`): `runComposeDeploy`'s failure branch runs `compose down` (volumes kept) before the record reads not-deployed. | **CLOSED 2026-09-23 — PROVEN-LIVE (`audits/r649-2026-09-23/`: outline with a never-healthy redis — 0 containers left, 3 volumes kept)** | `git show HEAD~1:documentation/backlog/OPEN-ITEMS.md` |
|
||||
|
||||
|
||||
@@ -806,7 +806,6 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-643** | **[P2-MEDIUM] The ruled chain leaves the automatic update leg AT MOST 15 MINUTES a night.** FOUND 2026-09-23 while writing the build plan for `09` §3 decision 11 (*updates after the off-site copy, before the full-system backup*). The off-site leg starts at W+105m (`cmd/controller/main.go:961`) and the full-system backup's gate opens at W+2h (`quiesce/quiesce.go:656`, span to W+6h); the legs are clock-scheduled, not chained. One step takes ~1 min when it works and ~2–6 min when it fails and is undone. Options and the recommendation (the full-system backup waits for the leg inside its own window; the leg stops starting steps at W+5h) are in `09` §6.4. **-- RULED 2026-09-23 (`09` §3 decision 20):** the full-system backup waits for the update leg inside its own window; the leg stops starting new steps at W+5h. Built with `09` §6.4 part 7. | **RULED — build with §6.4 part 7; owner: CC** |
|
||||
| **R-644** | **[P3-LOW] `gokapi` on scratch guest 9202 is crash-looping — 329 restarts by 2026-09-23 07:51 UTC, *password does not appear to be a SHA-1 hash* — and the controller still lists it deployed.** OBSERVED at the start of the 2026-09-23 session, not caused by it. The twenty-eight walk's teardown (2026-09-22) removed a `gokapi` container left by R-633 by name; a `gokapi` is running again, recorded `deployed: true`. Not investigated (scope). Likely the R-633/R-634 shape — a restore-then-remove race leaving a record — and a scratch-box fact, not a customer one; filed so the next drill does not read it as its own. | **OPEN — P3; owner: CC; investigate before the next drill on 9202** |
|
||||
| **R-645** | **[P3-LOW] Lifting an update hold by the operator CLI lets the recovery unit be re-captured with the FAILED new definition within seconds — the copy the hold sentence names is overwritten.** MEASURED 2026-09-23 on 9202 during the undo bake-off: docmost was held at 07:59:16Z after a failed 0.95.0 → 0.96.0 update; `--clear-restore-hold docmost` + the controller restart it requires ran at ~07:59:23Z, and at **07:59:26Z** the controller logged *Recovery unit captured for docmost* — the unit's `compose/docker-compose.yml` now named `docmost/docmost:0.96.0`, the version that had just failed. The hold sentence had pointed the household at that unit („saját meghajtó, … 09:55"). The hold is what keeps the nightly legs off a held app (`isHeld`, v0.238.1); once it is lifted by hand, the checksum-gated refresh sees a changed definition and captures it. **Who it hits:** an operator who lifts a hold to inspect or repair, before restoring. With the undo (R-637) a failed update no longer holds unless the undo also fails, so the path is rarer — it does not go away. Candidate shapes, none chosen: the CLI refuses to lift an UPDATE hold (only a restore lifts it); or the lift also puts the pin back; or the capture skips an app whose pin is not what it is running. Evidence: `audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt` (the invalid run) and README §"Three things". **-- 2026-09-23 (controller v0.263.0):** the undo never reads the unit, so this no longer affects the automatic undo; it still affects an operator who lifts a hold by hand before restoring. | **OPEN — P3; owner: CC** |
|
||||
| **R-649** | **[P2-MEDIUM] OPERATOR QUESTION: when `docker compose up` fails for the deploy's OWN reasons and leaves containers behind, what should the record say?** Found while diagnosing R-634 (2026-09-23). The backup race is fixed (v0.265.0), but `runComposeDeploy`'s failure branch (`stacks/deploy.go:421-435`) still writes `Deployed=false` without asking whether containers exist — e.g. a dependency's healthcheck that times out after its containers started. Since v0.262.0 the household can REMOVE such an app (`halfStateEvidence`), so nothing is stranded; but the page says „not installed" over containers that may be running. **Options:** (a) the failed deploy runs `compose down` on what it started, so „failed" means nothing runs — cost: a slow-but-healthy start is thrown away, and the household must press Deploy again; (b) keep the containers and record the deploy as FAILED-WITH-CONTAINERS, a state the page shows with a Remove button and the compose error in both languages — cost: a new state every surface must learn; (c) leave it (today). **Recommendation: (a)** — it keeps the R-634 promise (never containers under „not deployed") with the least new surface, and a deploy that failed is re-pressed anyway. Not taken unattended: it changes what a failed deploy does to what the household started (rule 1). | **OPEN — P2; owner: operator (decision), then CC** |
|
||||
| **R-650** | **[P3-LOW] A controller unit test that falls through to the REAL docker acts on the build host — DooPlex.** MEASURED 2026-09-23: my own first draft of `TestR634_VolumeLegNeverStopsADeployingApp` drove the real `DumpAppVolumesSafe` for its control case, and `docker run … tar` CREATED an empty volume `outline_outline_data` on DooPlex (removed by name, verified empty and unused; nothing else touched). The same draft of the stacks test ran a real `docker compose down` from a temp dir named `nextcloud` (no such project existed). Both tests now use seams (`dumpVolumesSafe`; `composeCmd` + an empty PATH). **The class is unguarded:** any test that reaches `exec.Command("docker", …)` on DooPlex acts on production Docker. Fix shape: a test-binary guard in `composeExec`/`execCommand` that refuses a real docker exec when running under `go test` unless an explicit env opt-in is set, plus a decoy test that proves it refuses. | **OPEN — P3; owner: CC (controller)** |
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
Reference in New Issue
Block a user