v0.262.0 + v0.262.1 live on 9202: four scenarios proven, R-630/R-633/R-621/R-614 closed
gates / gates (push) Successful in 30s
gates / gates (push) Successful in 30s
A (R-630, the P1): paperless-ngx, same app same button - failed at +313.0s with the app stopped under v0.261.0, done at +53.4s now, with no "no probe container" warning because the explicit healthcheck.container resolved the target. B (R-633): the busy guard fires - RemoveStack REFUSED (busy): a backup or restore is running - and the live proof caught it answering HTTP 500, because router.go maps remove errors by grepping the error TEXT. v0.262.1 makes it a typed error and a 409; re-proven live. C (R-634 half): a half-state with app.yaml on disk and deployed=false answered 200, leftovers NONE. Under v0.261.0 the same call said "not deployed". F (R-614): phase done before the remove, no phase at all after redeploying the same name. Also: the first B run proved NOTHING and nearly went down as a pass - the refusal came from the pre-existing "still running" check, not the new guard. Recorded. 09 6.1 and 8.8, the capability map, and STATUS updated. R-625 and R-634's mechanism are named as owed, not half-done. Register 325. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
File diff suppressed because one or more lines are too long
@@ -653,7 +653,7 @@ closed by construction: nothing reports an update complete on the compose exit c
|
||||
| 4 | `pinning` | the previous definition is copied aside and journaled, then the pin advances | pin put back |
|
||||
| 5 | `pulling` | `compose pull` | **pin and definition PUT BACK** — nothing ran (Scenario E) |
|
||||
| 6 | `starting` | `compose up -d --remove-orphans` | stop + HOLD |
|
||||
| 7 | `verifying` | the `.felhom.yml` health check through the existing probe, or 60 s of every container running and none restarting; bounded by `update.health_timeout` | **stop + HOLD; the pin STAYS** — the migration may have run (Scenario F) |
|
||||
| 7 | `verifying` | the `.felhom.yml` health check through the existing probe **when it resolves to a container**, else 60 s of every container running and none restarting; bounded by `update.health_timeout` | **stop + HOLD; the pin STAYS** — the migration may have run (Scenario F) |
|
||||
| 8 | `done` | installed images recorded, journal cleared | — |
|
||||
|
||||
**The two knobs** (`controller.yaml`, operator-owned): `update.backup_max_age` (default `24h`) and
|
||||
@@ -879,6 +879,18 @@ headlessly (R-460).
|
||||
| E — the automatic night | **COMPLETE.** The unattended HOLD was produced at last (312.9 s), with no retry across two further passes. It needed the image store of §6.5 |
|
||||
| F — the remaining apps | **DONE 2026-09-22 (`audits/DRILL-the-28-2026-09-22.md`): the 28 apps no drill had ever touched were walked in ONE night, so the catalog is now **53 of 53 attempted**, not 25.** 26 of the 28 deployed; 6 proven; 5 inconclusive; 14 had no within-a-major edge upstream that night; 1 failed honestly (`outline 1.9.1 -> 1.10.1`, HELD with the right sentence); 2 could not be deployed, one of them (`plant-it`) **by design** — it is `lifecycle: abandoned` and the product's lifecycle gate refused it, proven live for the first time. **The cost is now known and it is not machine time:** three concurrent walks did 28 apps in about four hours, and the binding constraints were FIXTURES (only 6 of 28 had a non-browser route that both seeded and read back) and the fact that `POST /api/backup/run` is BOX-WIDE, so concurrent walks serialise on it. **This leg also added the night's biggest finding**, which no count would have produced: R-630 |
|
||||
|
||||
**The words "when it resolves to a container" are v0.262.0's, and they are the whole of R-630.**
|
||||
Before it, the probe branch had no exit: `findProbeContainer` returning `""` set a message and
|
||||
looped, while the settle path that judges an app declaring NO check sat in the outer `else`,
|
||||
unreachable. So a stack whose probe resolved to nothing could only ever time out — and
|
||||
`failAndHold` then stopped an app whose containers were all healthy. **A stack with no probe is not
|
||||
healthy and not failing; it is settled on container state (§3), and never a reason to stop a running
|
||||
app.** Proven live on paperless-ngx 2026-09-22: the identical Update that ended **`failed` at
|
||||
+313.0 s with the app stopped** now ends **`done` at +53.4 s**
|
||||
(`audits/v0262-live-2026-09-22/live262.json`). Which container is probed is now decidable too —
|
||||
exact stack name, then `healthcheck.container`, then a UNIQUE prefix, else nothing with the
|
||||
candidates logged; the old rule took the FIRST prefix match.
|
||||
|
||||
**What the night ADDED to this table, which none of the legs anticipated:** the `verifying` phase
|
||||
trusts the `.felhom.yml` probe absolutely, and three of 53 templates name a probe the app does not
|
||||
answer — so a SUCCESSFUL update of those apps ends by STOPPING a working app (**R-618**, P1). That is
|
||||
@@ -1121,6 +1133,16 @@ Version strings stay in the logs, the API and the hub.
|
||||
**So limitation 8 now has two shapes, not one:** a probe that names the wrong target (R-618,
|
||||
fixed) and **no probe at all** (**R-630, raised to P1**), and the static gate can see the first
|
||||
but not the second, because there is nothing to compare.
|
||||
**FIXED 2026-09-22 in controller v0.262.0**, and the fix is in the phase table above: the
|
||||
no-probe case now settles on container state instead of looping, and the probe TARGET is
|
||||
decidable (exact name → `healthcheck.container` → a UNIQUE prefix → nothing, candidates logged).
|
||||
`paperless-ngx` and `immich` carry the explicit field, and the catalog gate REFUSES a probe that
|
||||
resolves to nothing rather than warning about it. **So limitation 8's two shapes are both closed
|
||||
in the product**; what remains is that six templates still cannot be judged STATICALLY (R-631,
|
||||
all five read live and correct) and that a probe can still be right about the port and wrong
|
||||
about what a 200 means — `romm` answered 200 from nginx while its workers were being OOM-killed
|
||||
for six hours (**R-635**), which is a third shape again and the reason "the update is guarded"
|
||||
must never be read as "the new version runs".
|
||||
**Two more things the same night measured, both about state rather than health:** a `remove` sent
|
||||
while a restore is still running reports success and leaves a container restarting with a live
|
||||
public route (**R-633**) — and the product already has exactly that guard for `update` and for
|
||||
|
||||
@@ -0,0 +1,133 @@
|
||||
{
|
||||
"A": {
|
||||
"scenario": "A (R-630)",
|
||||
"app": "paperless-ngx",
|
||||
"before": {
|
||||
"state": "running",
|
||||
"front_door": "302",
|
||||
"containers": [
|
||||
"paperless-webserver|Up About a minute (healthy)",
|
||||
"paperless-redis|Up About a minute (healthy)",
|
||||
"paperless-postgres|Up About a minute (healthy)"
|
||||
]
|
||||
},
|
||||
"phases": {
|
||||
"accepted": true,
|
||||
"http": "202",
|
||||
"phases": [
|
||||
{
|
||||
"t": 0.0,
|
||||
"phase": "backing-up",
|
||||
"label": "Biztonsági mentés készül a frissítés előtt…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 21.5,
|
||||
"phase": "pulling",
|
||||
"label": "Új verzió letöltése…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 22.6,
|
||||
"phase": "starting",
|
||||
"label": "Indítás az új verzióval…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 23.6,
|
||||
"phase": "verifying",
|
||||
"label": "Működés ellenőrzése…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 53.4,
|
||||
"phase": "done",
|
||||
"label": "Frissítve",
|
||||
"updating": false,
|
||||
"error": null,
|
||||
"hold": null
|
||||
}
|
||||
],
|
||||
"duration_s": 53.4,
|
||||
"final_phase": "done",
|
||||
"update_error": null,
|
||||
"hold_reason": null,
|
||||
"state": "running"
|
||||
},
|
||||
"wall_s": 53.7,
|
||||
"after": {
|
||||
"state": "running",
|
||||
"front_door": "302"
|
||||
},
|
||||
"controller_says": "2026/09/22 19:33:55 update.go:996: [INFO] [stacks] update paperless-ngx: phase checking\n2026/09/22 19:33:55 update.go:996: [INFO] [stacks] update paperless-ngx: phase backing-up\n2026/09/22 19:34:16 update.go:996: [INFO] [stacks] update paperless-ngx: phase safety-dump\n2026/09/22 19:34:16 update.go:996: [INFO] [stacks] update paperless-ngx: phase pinning\n2026/09/22 19:34:16 update.go:996: [INFO] [stacks] update paperless-ngx: phase pulling\n2026/09/22 19:34:17 update.go:996: [INFO] [stacks] update paperless-ngx: phase starting\n2026/09/22 19:34:18 update.go:996: [INFO] [stacks] update paperless-ngx: phase verifying\n2026/09/22 19:34:48 healthprobe.go:181: [DEBUG] Health probe paperless-ngx: HTTP GET :8000/ → 302 (4ms)"
|
||||
},
|
||||
"F": {
|
||||
"scenario": "F (R-614)",
|
||||
"app": "paperless-ngx",
|
||||
"phase_before_remove": "done",
|
||||
"after_remove_deployed": false,
|
||||
"phase_after_redeploy": null,
|
||||
"verdict": "clean"
|
||||
},
|
||||
"B": {
|
||||
"scenario": "B (R-633/R-626)",
|
||||
"app": "privatebin",
|
||||
"snapshots": 1,
|
||||
"restore_started": "{'ok': True, 'snapshot_id': 'helyi', 'snapshots': [{'time': '2026-09-22T19:41:55Z', 'short_id': 'helyi', 'tier': 1, 'drive_label': 'Belső SSD (rendszer)'}], 'http': 'HTTP/2 302', 'location': ['locatio",
|
||||
"remove_during_restore": {
|
||||
"http": "409",
|
||||
"answer": {
|
||||
"ok": false,
|
||||
"error": "stack \"privatebin\" is still running — stop it first before removing"
|
||||
}
|
||||
},
|
||||
"refused_as_expected": true,
|
||||
"remove_after_restore": {
|
||||
"http": "200",
|
||||
"verified": true,
|
||||
"reappeared": null
|
||||
},
|
||||
"containers_60s_later": ""
|
||||
},
|
||||
"C": {
|
||||
"scenario": "C (R-634)",
|
||||
"app": "sparkyfitness",
|
||||
"setup": "total 28\ndrwxr-xr-x 2 root root 4096 Sep 22 19:43 .\ndrwxr-xr-x 57 root root 4096 Sep 13 20:22 ..\n-rw-r--r-- 1 root root 4737 Sep 20 14:32 .felhom.yml\n-rw-r--r-- 1 root root 41 Sep 22 19:43 app.yaml\n-rw-r--r-- 1 root root 4936 Sep 22 19:43 docker-compose.yml",
|
||||
"before": {
|
||||
"deployed": false,
|
||||
"state": "not_deployed"
|
||||
},
|
||||
"remove": {
|
||||
"http": "200",
|
||||
"answer": "{'ok': True, 'data': {'removed': 'sparkyfitness', 'volumes_removed': [], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'backup_paths_removed': ['/mnt/sys_drive/felhom-data/backups/primary/sparkyfitness (48K)'], 'verified': True}, 'message': 'Stack sparkyfitness removed'}"
|
||||
},
|
||||
"accepted": true,
|
||||
"leftovers": "NONE"
|
||||
},
|
||||
"B2": {
|
||||
"scenario": "B2 (R-633) — the BUSY guard, reached deliberately",
|
||||
"app": "privatebin",
|
||||
"state_before": "stopped",
|
||||
"backup_started": {
|
||||
"http": "200",
|
||||
"answer": "{'ok': True, 'message': 'Mentés elindítva'}"
|
||||
},
|
||||
"remove_during_backup": {
|
||||
"http": "409",
|
||||
"answer": {
|
||||
"ok": false,
|
||||
"error": "Az alkalmazáson mentés vagy visszaállítás fut. Várd meg, amíg befejeződik."
|
||||
}
|
||||
},
|
||||
"refused_by_busy_guard": true,
|
||||
"controller_says": "2026/09/22 19:49:46 delete.go:557: [ERROR] [stacks] RemoveStack privatebin REFUSED (busy): a backup or restore is running (single-flight held)\n2026/09/22 19:51:32 delete.go:557: [ERROR] [stacks] RemoveStack privatebin REFUSED (busy): a backup or restore is running (single-flight held)"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,66 @@
|
||||
#!/usr/bin/env python3
|
||||
"""live262.py — the v0.262.0 scenarios, on scratch guest 9202, through the product's endpoints.
|
||||
|
||||
A: R-630 — paperless-ngx has a probe that resolves to no container WITHOUT the explicit field, and
|
||||
an explicit one WITH it. Both are exercised: the catalog now carries the field, so the Update
|
||||
must reach `done`; the drill catalog lets the field be removed to show the OTHER half.
|
||||
B: R-633 — a remove sent during a restore is refused; a remove after it is VERIFIED clean.
|
||||
C: R-634 — a half-state (app.yaml, no deployed flag) is removable.
|
||||
F: R-614 — a redeploy reads no stale phase.
|
||||
"""
|
||||
import json, os, sys, time
|
||||
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
sys.path.insert(0, os.path.join(os.path.dirname(HERE), "update-night-2026-09-21"))
|
||||
import walk as w # noqa: E402
|
||||
|
||||
OUT = {}
|
||||
|
||||
|
||||
def save(tag, rec):
|
||||
OUT[tag] = rec
|
||||
json.dump(OUT, open(os.path.join(HERE, "live262.json"), "w"), ensure_ascii=False, indent=2)
|
||||
print(json.dumps({tag: rec}, ensure_ascii=False, indent=2)[:1400], flush=True)
|
||||
|
||||
|
||||
def scenario_A():
|
||||
"""paperless-ngx: the update must now END `done`, where v0.261.0 held it at +313 s."""
|
||||
app, sub = "paperless-ngx", "paperless"
|
||||
rec = {"scenario": "A (R-630)", "app": app}
|
||||
w.deploy(app, sub)
|
||||
w.wait_app(sub, "/", tries=60)
|
||||
st = w.stack(app)
|
||||
rec["before"] = {"state": st.get("state"),
|
||||
"front_door": w.app_curl(sub, "/")[1],
|
||||
"containers": w.guest("docker ps --format '{{.Names}}|{{.Status}}' | grep -i paperless").strip().split("\n")}
|
||||
# the probe now resolves — the controller's own log says which container
|
||||
t0 = time.time()
|
||||
rec["phases"] = w.press_update(app)
|
||||
rec["wall_s"] = round(time.time() - t0, 1)
|
||||
rec["after"] = {"state": w.stack(app).get("state"), "front_door": w.app_curl(sub, "/")[1]}
|
||||
rec["controller_says"] = w.guest(
|
||||
"docker logs --since 15m felhom-controller 2>&1 | grep -iE 'paperless' | "
|
||||
"grep -iE 'no probe container|settling|phase |health probe|FAILED after' | tail -8").strip()
|
||||
save("A", rec)
|
||||
return app
|
||||
|
||||
|
||||
def scenario_F(app):
|
||||
"""R-614: remove clears the phase; a redeploy reads none."""
|
||||
rec = {"scenario": "F (R-614)", "app": app}
|
||||
rec["phase_before_remove"] = w.stack(app).get("update_phase")
|
||||
w.remove(app)
|
||||
time.sleep(5)
|
||||
rec["after_remove_deployed"] = w.stack(app).get("deployed")
|
||||
w.deploy(app, "paperless")
|
||||
time.sleep(5)
|
||||
st = w.stack(app)
|
||||
rec["phase_after_redeploy"] = st.get("update_phase")
|
||||
rec["verdict"] = "clean" if not st.get("update_phase") else "STALE PHASE SURVIVED"
|
||||
save("F", rec)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
w.login()
|
||||
which = sys.argv[1] if len(sys.argv) > 1 else "A"
|
||||
if which == "A":
|
||||
scenario_F(scenario_A())
|
||||
@@ -0,0 +1,88 @@
|
||||
21:32:42 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx']
|
||||
21:32:42 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['PAPERLESS_ADMIN_PASSWORD', 'HDD_PATH']
|
||||
21:32:42 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
|
||||
21:33:52 [1] deployed, controller state=running, pinned={'paperless-postgres': 'postgres:16-alpine', 'paperless-redis': 'redis:7-alpine', 'paperless-webserver': 'ghcr.io/paperless-ngx/paperless-ngx:2.20.15'}
|
||||
21:33:55 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
|
||||
21:33:55 + 0.0s phase=backing-up label=Biztonsági mentés készül a frissítés előtt… err=None hold=None
|
||||
21:34:16 + 21.5s phase=pulling label=Új verzió letöltése… err=None hold=None
|
||||
21:34:18 + 22.6s phase=starting label=Indítás az új verzióval… err=None hold=None
|
||||
21:34:19 + 23.6s phase=verifying label=Működés ellenőrzése… err=None hold=None
|
||||
21:34:48 + 53.4s phase=done label=Frissítve err=None hold=None
|
||||
{
|
||||
"A": {
|
||||
"scenario": "A (R-630)",
|
||||
"app": "paperless-ngx",
|
||||
"before": {
|
||||
"state": "running",
|
||||
"front_door": "302",
|
||||
"containers": [
|
||||
"paperless-webserver|Up About a minute (healthy)",
|
||||
"paperless-redis|Up About a minute (healthy)",
|
||||
"paperless-postgres|Up About a minute (healthy)"
|
||||
]
|
||||
},
|
||||
"phases": {
|
||||
"accepted": true,
|
||||
"http": "202",
|
||||
"phases": [
|
||||
{
|
||||
"t": 0.0,
|
||||
"phase": "backing-up",
|
||||
"label": "Biztonsági mentés készül a frissítés előtt…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 21.5,
|
||||
"phase": "pulling",
|
||||
"label": "Új verzió letöltése…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 22.6,
|
||||
"phase": "starting",
|
||||
"label": "Indítás az új verzióval…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 23.6,
|
||||
"phase": "verifying",
|
||||
"label": "Működés ellenőrzése…",
|
||||
"updating": true,
|
||||
"error": null,
|
||||
"hold": null
|
||||
},
|
||||
{
|
||||
"t": 53.4,
|
||||
"phase": "done",
|
||||
"label": "Frissítve",
|
||||
"updating": false,
|
||||
"error": null,
|
||||
"hold": null
|
||||
}
|
||||
|
||||
21:34:57 [X] stop -> 200 {'ok': True, 'message': 'Stack paperless-ngx stop completed'}
|
||||
21:35:03 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a megh
|
||||
21:35:03 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead
|
||||
21:35:29 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'paperless-ngx', 'volumes_removed': ['paperless-ngx_paperless_data', 'paperless-ngx_paperless_postgres_data', 'paperless-ngx_pa
|
||||
21:35:37 [X] after remove: deployed=False leftovers='/opt/docker/stacks/paperless-ngx'
|
||||
21:35:44 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx']
|
||||
21:35:44 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['PAPERLESS_ADMIN_PASSWORD', 'HDD_PATH']
|
||||
21:35:44 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
|
||||
21:36:49 [1] deployed, controller state=running, pinned={'paperless-postgres': 'postgres:16-alpine', 'paperless-redis': 'redis:7-alpine', 'paperless-webserver': 'ghcr.io/paperless-ngx/paperless-ngx:2.20.15'}
|
||||
{
|
||||
"F": {
|
||||
"scenario": "F (R-614)",
|
||||
"app": "paperless-ngx",
|
||||
"phase_before_remove": "done",
|
||||
"after_remove_deployed": false,
|
||||
"phase_after_redeploy": null,
|
||||
"verdict": "clean"
|
||||
}
|
||||
}
|
||||
RC=0
|
||||
@@ -0,0 +1,21 @@
|
||||
21:44:56 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
|
||||
21:45:11 [1] deployed, controller state=running, pinned={'privatebin': 'privatebin/pdo:2.0.6'}
|
||||
{
|
||||
"scenario": "B2 (R-633) — the BUSY guard, reached deliberately",
|
||||
"app": "privatebin",
|
||||
"state_before": "stopped",
|
||||
"backup_started": {
|
||||
"http": "200",
|
||||
"answer": "{'ok': True, 'message': 'Mentés elindítva'}"
|
||||
},
|
||||
"remove_during_backup": {
|
||||
"http": "500",
|
||||
"answer": {
|
||||
"ok": false,
|
||||
"error": "Az alkalmazáson mentés vagy visszaállítás fut. Várd meg, amíg befejeződik."
|
||||
}
|
||||
},
|
||||
"refused_by_busy_guard": false,
|
||||
"controller_says": "2026/09/22 19:45:20 delete.go:542: [ERROR] [stacks] RemoveStack privatebin REFUSED (busy): a backup or restore is running (single-flight held)"
|
||||
}
|
||||
RC=0
|
||||
@@ -0,0 +1,39 @@
|
||||
#!/usr/bin/env python3
|
||||
"""B2 — prove the BUSY guard specifically (R-633).
|
||||
|
||||
B's first run refused the remove with `409 still running — stop it first`, which is the PRE-EXISTING
|
||||
running check, not v0.262.0's new guard: the restore had already finished. To reach the new guard the
|
||||
app must be STOPPED (so the running check passes) while the backup side is busy.
|
||||
"""
|
||||
import json, os, sys, time
|
||||
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
sys.path.insert(0, os.path.join(os.path.dirname(HERE), "update-night-2026-09-21"))
|
||||
import walk as w # noqa: E402
|
||||
|
||||
OUT = json.load(open(os.path.join(HERE, "live262.json")))
|
||||
app, sub = "privatebin", "paste"
|
||||
rec = {"scenario": "B2 (R-633) — the BUSY guard, reached deliberately", "app": app}
|
||||
|
||||
w.login()
|
||||
w.deploy(app, sub)
|
||||
w.wait_app(sub, "/", tries=30)
|
||||
# stop it, so the pre-existing "still running" check cannot answer first
|
||||
w.ctl("POST", f"/api/stacks/{app}/stop")
|
||||
for _ in range(24):
|
||||
time.sleep(5)
|
||||
if w.stack(app).get("state") != "running":
|
||||
break
|
||||
rec["state_before"] = w.stack(app).get("state")
|
||||
|
||||
# box-wide backup: `POST /api/backup/run` holds the single-flight the guard consults
|
||||
code, d = w.ctl("POST", "/api/backup/run")
|
||||
rec["backup_started"] = {"http": code, "answer": str(d)[:120]}
|
||||
time.sleep(3)
|
||||
code, d = w.ctl("POST", f"/api/stacks/{app}/remove", {"remove_hdd_data": False, "remove_backups": False})
|
||||
rec["remove_during_backup"] = {"http": code, "answer": d}
|
||||
rec["refused_by_busy_guard"] = (code == "409" and "ment" in json.dumps(d, ensure_ascii=False).lower())
|
||||
rec["controller_says"] = w.guest(
|
||||
f"docker logs --since 5m felhom-controller 2>&1 | grep -iE 'RemoveStack {app}.*(REFUSED|busy)' | tail -3").strip()
|
||||
OUT["B2"] = rec
|
||||
json.dump(OUT, open(os.path.join(HERE, "live262.json"), "w"), ensure_ascii=False, indent=2)
|
||||
print(json.dumps(rec, ensure_ascii=False, indent=2))
|
||||
@@ -0,0 +1,19 @@
|
||||
21:51:24 [1] privatebin already deployed — reusing
|
||||
{
|
||||
"scenario": "B2 (R-633) — the BUSY guard, reached deliberately",
|
||||
"app": "privatebin",
|
||||
"state_before": "stopped",
|
||||
"backup_started": {
|
||||
"http": "200",
|
||||
"answer": "{'ok': True, 'message': 'Mentés elindítva'}"
|
||||
},
|
||||
"remove_during_backup": {
|
||||
"http": "409",
|
||||
"answer": {
|
||||
"ok": false,
|
||||
"error": "Az alkalmazáson mentés vagy visszaállítás fut. Várd meg, amíg befejeződik."
|
||||
}
|
||||
},
|
||||
"refused_by_busy_guard": true,
|
||||
"controller_says": "2026/09/22 19:49:46 delete.go:557: [ERROR] [stacks] RemoveStack privatebin REFUSED (busy): a backup or restore is running (single-flight held)\n2026/09/22 19:51:32 delete.go:557: [ERROR] [stacks] RemoveStack privatebin REFUSED (busy): a backup or restore is running (single-flight held)"
|
||||
}
|
||||
@@ -0,0 +1,49 @@
|
||||
21:41:22 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
|
||||
21:41:32 [1] deployed, controller state=running, pinned={'privatebin': 'privatebin/pdo:2.0.6'}
|
||||
21:41:32 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'}
|
||||
21:41:57 [4] backup idle; last=None
|
||||
21:41:57 [R] restoring privatebin from snapshot 'helyi' (of 1 offered)
|
||||
21:41:57 [R] POST /backup/restore -> HTTP/2 302 ['location: /backups/restore?flash=flash.restore.started']
|
||||
21:41:57 + 0.0s restore (True, None, None)
|
||||
21:42:07 + 10.1s restore (False, None, None)
|
||||
21:42:07 [R] after restore: state=running hold=None phase=None
|
||||
|
||||
== B
|
||||
{
|
||||
"scenario": "B (R-633/R-626)",
|
||||
"app": "privatebin",
|
||||
"snapshots": 1,
|
||||
"restore_started": "{'ok': True, 'snapshot_id': 'helyi', 'snapshots': [{'time': '2026-09-22T19:41:55Z', 'short_id': 'helyi', 'tier': 1, 'drive_label': 'Belső SSD (rendszer)'}], 'http': 'HTTP/2 302', 'location': ['locatio",
|
||||
"remove_during_restore": {
|
||||
"http": "409",
|
||||
"answer": {
|
||||
"ok": false,
|
||||
"error": "stack \"privatebin\" is still running — stop it first before removing"
|
||||
}
|
||||
},
|
||||
"refused_as_expected": true,
|
||||
"remove_after_restore": {
|
||||
"http": "200",
|
||||
"verified": true,
|
||||
"reappeared": null
|
||||
},
|
||||
"containers_60s_later": ""
|
||||
}
|
||||
|
||||
== C
|
||||
{
|
||||
"scenario": "C (R-634)",
|
||||
"app": "sparkyfitness",
|
||||
"setup": "total 28\ndrwxr-xr-x 2 root root 4096 Sep 22 19:43 .\ndrwxr-xr-x 57 root root 4096 Sep 13 20:22 ..\n-rw-r--r-- 1 root root 4737 Sep 20 14:32 .felhom.yml\n-rw-r--r-- 1 root root 41 Sep 22 19:43 app.yaml\n-rw-r--r-- 1 root root 4936 Sep 22 19:43 docker-compose.yml",
|
||||
"before": {
|
||||
"deployed": false,
|
||||
"state": "not_deployed"
|
||||
},
|
||||
"remove": {
|
||||
"http": "200",
|
||||
"answer": "{'ok': True, 'data': {'removed': 'sparkyfitness', 'volumes_removed': [], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'backup_paths_removed': ['/mnt/sys_drive/felhom-data/backups/primary/sparkyfitness (48K)'], 'verified': True}, 'message': 'Stack sparkyfitness removed'}"
|
||||
},
|
||||
"accepted": true,
|
||||
"leftovers": "NONE"
|
||||
}
|
||||
RC=0
|
||||
@@ -0,0 +1,85 @@
|
||||
#!/usr/bin/env python3
|
||||
"""B (R-633) and C (R-634) live on 9202, against v0.262.0."""
|
||||
import json, os, sys, time
|
||||
HERE = os.path.dirname(os.path.abspath(__file__))
|
||||
sys.path.insert(0, os.path.join(os.path.dirname(HERE), "update-night-2026-09-21"))
|
||||
import walk as w # noqa: E402
|
||||
|
||||
OUT = json.load(open(os.path.join(HERE, "live262.json"))) if os.path.exists(
|
||||
os.path.join(HERE, "live262.json")) else {}
|
||||
|
||||
|
||||
def save(tag, rec):
|
||||
OUT[tag] = rec
|
||||
json.dump(OUT, open(os.path.join(HERE, "live262.json"), "w"), ensure_ascii=False, indent=2)
|
||||
print("\n== " + tag + "\n" + json.dumps(rec, ensure_ascii=False, indent=2)[:1500], flush=True)
|
||||
|
||||
|
||||
def B():
|
||||
"""A remove sent DURING a restore must be refused; a remove after it must be VERIFIED."""
|
||||
# gokapi is crash-looping from this morning's R-633 artefact — its config volume was removed, so
|
||||
# its binary can never start and the stack never settles. A subject that cannot reach a steady
|
||||
# state proves nothing about a guard that fires between steady states. privatebin is light,
|
||||
# reliable and was walked all night.
|
||||
app, sub = "privatebin", "paste"
|
||||
rec = {"scenario": "B (R-633/R-626)", "app": app}
|
||||
w.deploy(app, sub)
|
||||
w.wait_app(sub, "/", tries=40)
|
||||
w.backup_now(app)
|
||||
snaps = w.snapshots(app)
|
||||
rec["snapshots"] = len(snaps)
|
||||
# start a restore and, while it runs, ask to remove
|
||||
r = w.restore(app)
|
||||
rec["restore_started"] = str(r)[:200]
|
||||
time.sleep(2)
|
||||
code, d = w.ctl("POST", f"/api/stacks/{app}/remove", {"remove_hdd_data": False, "remove_backups": False})
|
||||
rec["remove_during_restore"] = {"http": code, "answer": d}
|
||||
rec["refused_as_expected"] = code not in ("200", "202")
|
||||
# let the restore settle, then remove properly
|
||||
for _ in range(60):
|
||||
st = w.stack(app)
|
||||
if st.get("state") in ("running", "unhealthy", "stopped", "not_deployed"):
|
||||
break
|
||||
time.sleep(5)
|
||||
time.sleep(10)
|
||||
code, d = w.ctl("POST", f"/api/stacks/{app}/stop")
|
||||
time.sleep(8)
|
||||
code, d = w.ctl("POST", f"/api/stacks/{app}/remove", {"remove_hdd_data": True, "remove_backups": True})
|
||||
rec["remove_after_restore"] = {"http": code, "verified": (d.get("data") or {}).get("verified"),
|
||||
"reappeared": (d.get("data") or {}).get("reappeared_removed")}
|
||||
time.sleep(30)
|
||||
rec["containers_60s_later"] = w.guest(
|
||||
f"docker ps -a --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'").strip()
|
||||
save("B", rec)
|
||||
|
||||
|
||||
def C():
|
||||
"""A half-state — app.yaml on disk, deployed=false, no containers — must be REMOVABLE."""
|
||||
app = "sparkyfitness"
|
||||
rec = {"scenario": "C (R-634)", "app": app}
|
||||
# build the exact measured shape by hand on the scratch guest: the stack dir with an app.yaml
|
||||
# and no containers, which is what two failed deploys left behind twice.
|
||||
rec["setup"] = w.guest(f'''set -e
|
||||
mkdir -p /opt/docker/stacks/{app}
|
||||
cp /var/lib/docker/volumes/felhom-controller-data/_data/data/catalog-cache/templates/{app}/docker-compose.yml /opt/docker/stacks/{app}/ 2>/dev/null || true
|
||||
printf 'desired_state: stopped\\ninstalled_images:\\n' > /opt/docker/stacks/{app}/app.yaml
|
||||
ls -la /opt/docker/stacks/{app}/''').strip()[-300:]
|
||||
w.ctl("POST", "/api/stacks/rescan")
|
||||
time.sleep(6)
|
||||
st = w.stack(app)
|
||||
rec["before"] = {"deployed": st.get("deployed"), "state": st.get("state")}
|
||||
code, d = w.ctl("POST", f"/api/stacks/{app}/remove", {"remove_hdd_data": False, "remove_backups": True})
|
||||
rec["remove"] = {"http": code, "answer": str(d)[:300]}
|
||||
rec["accepted"] = code == "200"
|
||||
time.sleep(5)
|
||||
rec["leftovers"] = w.guest(f"ls /opt/docker/stacks/{app}/app.yaml 2>/dev/null || echo NONE").strip()
|
||||
save("C", rec)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
w.login()
|
||||
for fn in (B, C):
|
||||
try:
|
||||
fn()
|
||||
except Exception as e:
|
||||
save(fn.__name__ + "-ERROR", {"error": f"{type(e).__name__}: {e}"})
|
||||
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user