v0.262.0 + v0.262.1 live on 9202: four scenarios proven, R-630/R-633/R-621/R-614 closed
gates / gates (push) Successful in 30s

A (R-630, the P1): paperless-ngx, same app same button - failed at +313.0s with the app stopped
under v0.261.0, done at +53.4s now, with no "no probe container" warning because the explicit
healthcheck.container resolved the target.

B (R-633): the busy guard fires - RemoveStack REFUSED (busy): a backup or restore is running - and
the live proof caught it answering HTTP 500, because router.go maps remove errors by grepping the
error TEXT. v0.262.1 makes it a typed error and a 409; re-proven live.

C (R-634 half): a half-state with app.yaml on disk and deployed=false answered 200, leftovers NONE.
Under v0.261.0 the same call said "not deployed".

F (R-614): phase done before the remove, no phase at all after redeploying the same name.

Also: the first B run proved NOTHING and nearly went down as a pass - the refusal came from the
pre-existing "still running" check, not the new guard. Recorded.

09 6.1 and 8.8, the capability map, and STATUS updated. R-625 and R-634's mechanism are named as
owed, not half-done. Register 325.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-22 21:53:53 +02:00
parent 1de6aaf904
commit 22439b0e43
12 changed files with 544 additions and 39 deletions
+15 -32
View File
@@ -1,44 +1,27 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-09-22 (overnight) — I installed and tested all 28 apps that no test had ever touched. Every app in our catalogue has now been tried at least once. Three things are quietly wrong, and one of them stops a working app.**
**Updated 2026-09-22 (late) — I fixed the six faults the two drill nights found in the update, delete and hold machinery, and shipped the six app versions you approved. One thing needs your word: whether the fleet moves to the new controller.**
**Decisions I took on my own: none.**
**What I did.** Each of the twenty-eight got the same walk: install it at the version our catalogue offers today, put real data in through the app's own front door, back it up, update it if a newer version really exists, **restore it from that backup and read the data back again**, then delete it and check a minute later that nothing came back. That restore step is new — the update night skipped it. **Twenty-six of the twenty-eight installed. Six are proven end to end.** Fourteen had no newer version to move to tonight. One failed honestly. Two would not install, and one of those is meant not to.
**The one that mattered most is fixed and proven.** An app with no health check used to be **shut down by a successful update** — the machine waited five minutes for a check that could never arrive, then stopped a working app. Paperless-ngx, same app, same button: **before, it failed after 5 minutes and the app went dark. Now it finishes in 53 seconds and keeps running.**
**The thing I would fix first — an app with no health check gets shut down by a successful update.** Paperless-ngx is never health-checked at all: its containers are named differently from the app, so the machine looks for one, finds nothing, and moves on without a word. I always thought that was just a missing badge. It is not. **The update waits five minutes for a health check that can never arrive, then declares failure and shuts the working app down.** All three of its containers were healthy the whole time. The machine says so in its own words: *"not healthy within 5m0s (last: no probe container) — stopping and HOLDING the app"*. Every household running Paperless who presses Update loses their app and is sent to a restore they do not need.
**Five more, all proven on the test machine.**
- **Deleting an app while it is being backed up or restored is now refused**, with a plain sentence telling you to wait — instead of quietly tearing it down and leaving a ghost behind.
- **A delete now checks its own work.** The machine watches for 25 seconds afterwards and removes anything that comes back, and says whether it verified.
- **An app the machine has lost track of can now be deleted.** Before, if its record went wrong, no button worked and only a command line could clear it.
- **A failed update now keeps the app's own log** before shutting it down. Twice we lost the only evidence of why.
- **Deleting an app clears its old update status**, so a fresh install of the same app no longer shows a stale "Updated".
**Two more, both about the machine losing track of an app rather than its health.**
- **Deleting an app while it is being restored leaves a ghost.** Both buttons say they worked. The app vanishes from every screen, and a container keeps restarting on the machine, still holding a public web address. **The machine already knows how to refuse this** — it refuses an *update* while a backup runs, and refuses a second *restore* while one is going, and it even names which app is blocking. Delete has no such guard.
- **An app can be running perfectly while the machine records it as not installed — and then it cannot be deleted.** I saw this three times. Two only happened when several jobs ran at once; one happened on its own, repeatably. In that state there is no button that works.
**The six versions you approved are live on the catalogue** — Emby, Ghost, Immich, Radarr, Sonarr, Termix. **None of them is installed on either demo machine**, so nothing updated; they simply show as available.
**In all three cases I needed a command line to clean up what the product could not. A household has none.**
**What I did not do, and it is on purpose.** Two items from the plan are untouched and named rather than half-finished: finding out *why* an app's record goes wrong in the first place (I fixed the consequence, not the cause), and making a held app stop offering an Update button it will refuse.
**The best thing I saw.** Two apps keep their files outside the database, and their local copy does not hold those files. When I asked to restore them, the machine **refused** — and said, in plain Hungarian, that it will not put an old database on top of files it does not have, that the files stay where they are, and which button does work. That is exactly right.
**What went wrong on my side.** I lost **44 minutes** to my own progress-watchers: they waited for a build that had already succeeded, because each was watching for a name its own command contained. The same bug cost me a pile of stuck watchers earlier in the day. It is now written down as a rule so it does not happen a third time. I also nearly recorded one test as passing when it had proved nothing — the refusal I saw came from an older rule, not the new one. I caught it and re-ran it properly.
**What I got wrong.** My own test script had three bugs that cost nine apps their walk. I found them, fixed them, and walked those nine again one at a time — and that second pass is what corrected my conclusions and produced one of the six proofs. The report names all three.
**Rows opened and closed.** Four closed, one narrowed to what is still unknown. The list stands at 325.
**Rows opened and closed.** Three new (the two above, plus the one that raised Paperless to urgent). Two closed. The list went from 321 to 323.
**What needs you — one question.**
1. **Shall the fleet move to controller 0.262.1?** Right now only the test machine has it. **My recommendation is yes:** every change here only refuses, waits, records, or removes what someone already asked to remove — none of them makes the machine do more on its own. *If you do nothing:* both demo machines stay on 0.261.0 and keep all six faults, including the one that shuts down a working app. Peti's machine is parked and would take it only if it ever comes back online.
**What needs you.**
1. **The second promotion list** — six app versions this night proved safe enough to move on the real catalogue, and six named that must not move, each with the reason. It is in the report. Moving a version is your call, never mine. *If you do nothing:* nothing breaks; those apps drift further from upstream each month.
**Nothing on your own machine, the tester's machine, or the off-site box was touched. No product code was written. The real catalogue was never changed — I checked its version lines against the start of the night and not one differs.**
---
**Evening addition, 2026-09-22 — you heard the fans, and you were right.**
**RomM was cooking the HP box and it was my fault.** I moved it to a new version that morning. The update said it worked, and it did — for two hours. Then it ran out of memory and spent six hours killing and restarting its own workers, about seven times a minute, burning five processor cores.
**CORRECTION — the machine DID warn you, and I was wrong to say it did not.** You showed me the two e-mails: *"Alkalmazás memóriája elfogyott: romm"*, at 11:09 and again at 17:48. They are in the hub's Events and Notifications tabs too. **I wrote "nothing warned anyone" without opening either tab** — I went by an old note saying this signal was unproven and turned that into "it did not happen". That is the same mistake I made with the Hetzner tickets four days ago.
**What is actually wrong is smaller and real:** the machine sends **one** warning per app start. Six hours of trouble and 4,530 worker deaths produced **one e-mail** — the same e-mail a single harmless hiccup would send. It never gets louder, and the app keeps showing as running. The hub *did* have the full picture on the App Telemetry page (RomM: 5,023 errors, 632 warnings, while every other app showed zero), but nothing turns that into a second, louder alert. **That is why a correct warning still got missed, and it is now written down as its own item.**
**Fixed, and proven under real load.** Giving it more memory was not enough — I measured that rather than assuming it. The real cause was that the app starts **four web workers**, which is a server setting, on a box serving one household. It now starts **two**. I then drove **26,645 real requests** at it for five minutes: memory stayed between 416 and 614 MB against its 768 MB ceiling and **trended down**, with **zero** worker deaths. At rest it now uses **1.6% of a core** instead of 500%.
**Two things I got wrong on the way, both worth knowing.** My first memory number (768 MB) was a guess and it did not fix anything. And my first load test was pointed at the wrong web address — every request bounced off the front door in 9 milliseconds while the counter reported 14,026 successes. I caught that only because the app looked *too* idle. Both are written up.
**The lesson that outlives RomM.** Every "proven" result this week measured an app for the **minutes of the test**. RomM passed everything and broke two hours later. **Proven has meant "the update worked and the data survived", not "the new version runs".** That gap is now on the record.
**Still worth an eye:** RomM sits at 610 MB of its 768 MB, about 79%. It works, but it is not roomy, and nothing watches it.
**Nothing on your own machine or the off-site box was touched. The demo machines were not touched — they only see the six new version badges.**
File diff suppressed because one or more lines are too long
@@ -653,7 +653,7 @@ closed by construction: nothing reports an update complete on the compose exit c
| 4 | `pinning` | the previous definition is copied aside and journaled, then the pin advances | pin put back |
| 5 | `pulling` | `compose pull` | **pin and definition PUT BACK** — nothing ran (Scenario E) |
| 6 | `starting` | `compose up -d --remove-orphans` | stop + HOLD |
| 7 | `verifying` | the `.felhom.yml` health check through the existing probe, or 60 s of every container running and none restarting; bounded by `update.health_timeout` | **stop + HOLD; the pin STAYS** — the migration may have run (Scenario F) |
| 7 | `verifying` | the `.felhom.yml` health check through the existing probe **when it resolves to a container**, else 60 s of every container running and none restarting; bounded by `update.health_timeout` | **stop + HOLD; the pin STAYS** — the migration may have run (Scenario F) |
| 8 | `done` | installed images recorded, journal cleared | — |
**The two knobs** (`controller.yaml`, operator-owned): `update.backup_max_age` (default `24h`) and
@@ -879,6 +879,18 @@ headlessly (R-460).
| E — the automatic night | **COMPLETE.** The unattended HOLD was produced at last (312.9 s), with no retry across two further passes. It needed the image store of §6.5 |
| F — the remaining apps | **DONE 2026-09-22 (`audits/DRILL-the-28-2026-09-22.md`): the 28 apps no drill had ever touched were walked in ONE night, so the catalog is now **53 of 53 attempted**, not 25.** 26 of the 28 deployed; 6 proven; 5 inconclusive; 14 had no within-a-major edge upstream that night; 1 failed honestly (`outline 1.9.1 -> 1.10.1`, HELD with the right sentence); 2 could not be deployed, one of them (`plant-it`) **by design** — it is `lifecycle: abandoned` and the product's lifecycle gate refused it, proven live for the first time. **The cost is now known and it is not machine time:** three concurrent walks did 28 apps in about four hours, and the binding constraints were FIXTURES (only 6 of 28 had a non-browser route that both seeded and read back) and the fact that `POST /api/backup/run` is BOX-WIDE, so concurrent walks serialise on it. **This leg also added the night's biggest finding**, which no count would have produced: R-630 |
**The words "when it resolves to a container" are v0.262.0's, and they are the whole of R-630.**
Before it, the probe branch had no exit: `findProbeContainer` returning `""` set a message and
looped, while the settle path that judges an app declaring NO check sat in the outer `else`,
unreachable. So a stack whose probe resolved to nothing could only ever time out — and
`failAndHold` then stopped an app whose containers were all healthy. **A stack with no probe is not
healthy and not failing; it is settled on container state (§3), and never a reason to stop a running
app.** Proven live on paperless-ngx 2026-09-22: the identical Update that ended **`failed` at
+313.0 s with the app stopped** now ends **`done` at +53.4 s**
(`audits/v0262-live-2026-09-22/live262.json`). Which container is probed is now decidable too —
exact stack name, then `healthcheck.container`, then a UNIQUE prefix, else nothing with the
candidates logged; the old rule took the FIRST prefix match.
**What the night ADDED to this table, which none of the legs anticipated:** the `verifying` phase
trusts the `.felhom.yml` probe absolutely, and three of 53 templates name a probe the app does not
answer — so a SUCCESSFUL update of those apps ends by STOPPING a working app (**R-618**, P1). That is
@@ -1121,6 +1133,16 @@ Version strings stay in the logs, the API and the hub.
**So limitation 8 now has two shapes, not one:** a probe that names the wrong target (R-618,
fixed) and **no probe at all** (**R-630, raised to P1**), and the static gate can see the first
but not the second, because there is nothing to compare.
**FIXED 2026-09-22 in controller v0.262.0**, and the fix is in the phase table above: the
no-probe case now settles on container state instead of looping, and the probe TARGET is
decidable (exact name → `healthcheck.container` → a UNIQUE prefix → nothing, candidates logged).
`paperless-ngx` and `immich` carry the explicit field, and the catalog gate REFUSES a probe that
resolves to nothing rather than warning about it. **So limitation 8's two shapes are both closed
in the product**; what remains is that six templates still cannot be judged STATICALLY (R-631,
all five read live and correct) and that a probe can still be right about the port and wrong
about what a 200 means — `romm` answered 200 from nginx while its workers were being OOM-killed
for six hours (**R-635**), which is a third shape again and the reason "the update is guarded"
must never be read as "the new version runs".
**Two more things the same night measured, both about state rather than health:** a `remove` sent
while a restore is still running reports success and leaves a container restarting with a live
public route (**R-633**) — and the product already has exactly that guard for `update` and for
@@ -0,0 +1,133 @@
{
"A": {
"scenario": "A (R-630)",
"app": "paperless-ngx",
"before": {
"state": "running",
"front_door": "302",
"containers": [
"paperless-webserver|Up About a minute (healthy)",
"paperless-redis|Up About a minute (healthy)",
"paperless-postgres|Up About a minute (healthy)"
]
},
"phases": {
"accepted": true,
"http": "202",
"phases": [
{
"t": 0.0,
"phase": "backing-up",
"label": "Biztonsági mentés készül a frissítés előtt…",
"updating": true,
"error": null,
"hold": null
},
{
"t": 21.5,
"phase": "pulling",
"label": "Új verzió letöltése…",
"updating": true,
"error": null,
"hold": null
},
{
"t": 22.6,
"phase": "starting",
"label": "Indítás az új verzióval…",
"updating": true,
"error": null,
"hold": null
},
{
"t": 23.6,
"phase": "verifying",
"label": "Működés ellenőrzése…",
"updating": true,
"error": null,
"hold": null
},
{
"t": 53.4,
"phase": "done",
"label": "Frissítve",
"updating": false,
"error": null,
"hold": null
}
],
"duration_s": 53.4,
"final_phase": "done",
"update_error": null,
"hold_reason": null,
"state": "running"
},
"wall_s": 53.7,
"after": {
"state": "running",
"front_door": "302"
},
"controller_says": "2026/09/22 19:33:55 update.go:996: [INFO] [stacks] update paperless-ngx: phase checking\n2026/09/22 19:33:55 update.go:996: [INFO] [stacks] update paperless-ngx: phase backing-up\n2026/09/22 19:34:16 update.go:996: [INFO] [stacks] update paperless-ngx: phase safety-dump\n2026/09/22 19:34:16 update.go:996: [INFO] [stacks] update paperless-ngx: phase pinning\n2026/09/22 19:34:16 update.go:996: [INFO] [stacks] update paperless-ngx: phase pulling\n2026/09/22 19:34:17 update.go:996: [INFO] [stacks] update paperless-ngx: phase starting\n2026/09/22 19:34:18 update.go:996: [INFO] [stacks] update paperless-ngx: phase verifying\n2026/09/22 19:34:48 healthprobe.go:181: [DEBUG] Health probe paperless-ngx: HTTP GET :8000/ → 302 (4ms)"
},
"F": {
"scenario": "F (R-614)",
"app": "paperless-ngx",
"phase_before_remove": "done",
"after_remove_deployed": false,
"phase_after_redeploy": null,
"verdict": "clean"
},
"B": {
"scenario": "B (R-633/R-626)",
"app": "privatebin",
"snapshots": 1,
"restore_started": "{'ok': True, 'snapshot_id': 'helyi', 'snapshots': [{'time': '2026-09-22T19:41:55Z', 'short_id': 'helyi', 'tier': 1, 'drive_label': 'Belső SSD (rendszer)'}], 'http': 'HTTP/2 302', 'location': ['locatio",
"remove_during_restore": {
"http": "409",
"answer": {
"ok": false,
"error": "stack \"privatebin\" is still running — stop it first before removing"
}
},
"refused_as_expected": true,
"remove_after_restore": {
"http": "200",
"verified": true,
"reappeared": null
},
"containers_60s_later": ""
},
"C": {
"scenario": "C (R-634)",
"app": "sparkyfitness",
"setup": "total 28\ndrwxr-xr-x 2 root root 4096 Sep 22 19:43 .\ndrwxr-xr-x 57 root root 4096 Sep 13 20:22 ..\n-rw-r--r-- 1 root root 4737 Sep 20 14:32 .felhom.yml\n-rw-r--r-- 1 root root 41 Sep 22 19:43 app.yaml\n-rw-r--r-- 1 root root 4936 Sep 22 19:43 docker-compose.yml",
"before": {
"deployed": false,
"state": "not_deployed"
},
"remove": {
"http": "200",
"answer": "{'ok': True, 'data': {'removed': 'sparkyfitness', 'volumes_removed': [], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'backup_paths_removed': ['/mnt/sys_drive/felhom-data/backups/primary/sparkyfitness (48K)'], 'verified': True}, 'message': 'Stack sparkyfitness removed'}"
},
"accepted": true,
"leftovers": "NONE"
},
"B2": {
"scenario": "B2 (R-633) — the BUSY guard, reached deliberately",
"app": "privatebin",
"state_before": "stopped",
"backup_started": {
"http": "200",
"answer": "{'ok': True, 'message': 'Mentés elindítva'}"
},
"remove_during_backup": {
"http": "409",
"answer": {
"ok": false,
"error": "Az alkalmazáson mentés vagy visszaállítás fut. Várd meg, amíg befejeződik."
}
},
"refused_by_busy_guard": true,
"controller_says": "2026/09/22 19:49:46 delete.go:557: [ERROR] [stacks] RemoveStack privatebin REFUSED (busy): a backup or restore is running (single-flight held)\n2026/09/22 19:51:32 delete.go:557: [ERROR] [stacks] RemoveStack privatebin REFUSED (busy): a backup or restore is running (single-flight held)"
}
}
@@ -0,0 +1,66 @@
#!/usr/bin/env python3
"""live262.py — the v0.262.0 scenarios, on scratch guest 9202, through the product's endpoints.
A: R-630 — paperless-ngx has a probe that resolves to no container WITHOUT the explicit field, and
an explicit one WITH it. Both are exercised: the catalog now carries the field, so the Update
must reach `done`; the drill catalog lets the field be removed to show the OTHER half.
B: R-633 — a remove sent during a restore is refused; a remove after it is VERIFIED clean.
C: R-634 — a half-state (app.yaml, no deployed flag) is removable.
F: R-614 — a redeploy reads no stale phase.
"""
import json, os, sys, time
HERE = os.path.dirname(os.path.abspath(__file__))
sys.path.insert(0, os.path.join(os.path.dirname(HERE), "update-night-2026-09-21"))
import walk as w # noqa: E402
OUT = {}
def save(tag, rec):
OUT[tag] = rec
json.dump(OUT, open(os.path.join(HERE, "live262.json"), "w"), ensure_ascii=False, indent=2)
print(json.dumps({tag: rec}, ensure_ascii=False, indent=2)[:1400], flush=True)
def scenario_A():
"""paperless-ngx: the update must now END `done`, where v0.261.0 held it at +313 s."""
app, sub = "paperless-ngx", "paperless"
rec = {"scenario": "A (R-630)", "app": app}
w.deploy(app, sub)
w.wait_app(sub, "/", tries=60)
st = w.stack(app)
rec["before"] = {"state": st.get("state"),
"front_door": w.app_curl(sub, "/")[1],
"containers": w.guest("docker ps --format '{{.Names}}|{{.Status}}' | grep -i paperless").strip().split("\n")}
# the probe now resolves — the controller's own log says which container
t0 = time.time()
rec["phases"] = w.press_update(app)
rec["wall_s"] = round(time.time() - t0, 1)
rec["after"] = {"state": w.stack(app).get("state"), "front_door": w.app_curl(sub, "/")[1]}
rec["controller_says"] = w.guest(
"docker logs --since 15m felhom-controller 2>&1 | grep -iE 'paperless' | "
"grep -iE 'no probe container|settling|phase |health probe|FAILED after' | tail -8").strip()
save("A", rec)
return app
def scenario_F(app):
"""R-614: remove clears the phase; a redeploy reads none."""
rec = {"scenario": "F (R-614)", "app": app}
rec["phase_before_remove"] = w.stack(app).get("update_phase")
w.remove(app)
time.sleep(5)
rec["after_remove_deployed"] = w.stack(app).get("deployed")
w.deploy(app, "paperless")
time.sleep(5)
st = w.stack(app)
rec["phase_after_redeploy"] = st.get("update_phase")
rec["verdict"] = "clean" if not st.get("update_phase") else "STALE PHASE SURVIVED"
save("F", rec)
if __name__ == "__main__":
w.login()
which = sys.argv[1] if len(sys.argv) > 1 else "A"
if which == "A":
scenario_F(scenario_A())
@@ -0,0 +1,88 @@
21:32:42 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx']
21:32:42 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['PAPERLESS_ADMIN_PASSWORD', 'HDD_PATH']
21:32:42 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
21:33:52 [1] deployed, controller state=running, pinned={'paperless-postgres': 'postgres:16-alpine', 'paperless-redis': 'redis:7-alpine', 'paperless-webserver': 'ghcr.io/paperless-ngx/paperless-ngx:2.20.15'}
21:33:55 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
21:33:55 + 0.0s phase=backing-up label=Biztonsági mentés készül a frissítés előtt… err=None hold=None
21:34:16 + 21.5s phase=pulling label=Új verzió letöltése… err=None hold=None
21:34:18 + 22.6s phase=starting label=Indítás az új verzióval… err=None hold=None
21:34:19 + 23.6s phase=verifying label=Működés ellenőrzése… err=None hold=None
21:34:48 + 53.4s phase=done label=Frissítve err=None hold=None
{
"A": {
"scenario": "A (R-630)",
"app": "paperless-ngx",
"before": {
"state": "running",
"front_door": "302",
"containers": [
"paperless-webserver|Up About a minute (healthy)",
"paperless-redis|Up About a minute (healthy)",
"paperless-postgres|Up About a minute (healthy)"
]
},
"phases": {
"accepted": true,
"http": "202",
"phases": [
{
"t": 0.0,
"phase": "backing-up",
"label": "Biztonsági mentés készül a frissítés előtt…",
"updating": true,
"error": null,
"hold": null
},
{
"t": 21.5,
"phase": "pulling",
"label": "Új verzió letöltése…",
"updating": true,
"error": null,
"hold": null
},
{
"t": 22.6,
"phase": "starting",
"label": "Indítás az új verzióval…",
"updating": true,
"error": null,
"hold": null
},
{
"t": 23.6,
"phase": "verifying",
"label": "Működés ellenőrzése…",
"updating": true,
"error": null,
"hold": null
},
{
"t": 53.4,
"phase": "done",
"label": "Frissítve",
"updating": false,
"error": null,
"hold": null
}
21:34:57 [X] stop -> 200 {'ok': True, 'message': 'Stack paperless-ngx stop completed'}
21:35:03 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a megh
21:35:03 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead
21:35:29 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'paperless-ngx', 'volumes_removed': ['paperless-ngx_paperless_data', 'paperless-ngx_paperless_postgres_data', 'paperless-ngx_pa
21:35:37 [X] after remove: deployed=False leftovers='/opt/docker/stacks/paperless-ngx'
21:35:44 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx']
21:35:44 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['PAPERLESS_ADMIN_PASSWORD', 'HDD_PATH']
21:35:44 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
21:36:49 [1] deployed, controller state=running, pinned={'paperless-postgres': 'postgres:16-alpine', 'paperless-redis': 'redis:7-alpine', 'paperless-webserver': 'ghcr.io/paperless-ngx/paperless-ngx:2.20.15'}
{
"F": {
"scenario": "F (R-614)",
"app": "paperless-ngx",
"phase_before_remove": "done",
"after_remove_deployed": false,
"phase_after_redeploy": null,
"verdict": "clean"
}
}
RC=0
@@ -0,0 +1,21 @@
21:44:56 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
21:45:11 [1] deployed, controller state=running, pinned={'privatebin': 'privatebin/pdo:2.0.6'}
{
"scenario": "B2 (R-633) — the BUSY guard, reached deliberately",
"app": "privatebin",
"state_before": "stopped",
"backup_started": {
"http": "200",
"answer": "{'ok': True, 'message': 'Mentés elindítva'}"
},
"remove_during_backup": {
"http": "500",
"answer": {
"ok": false,
"error": "Az alkalmazáson mentés vagy visszaállítás fut. Várd meg, amíg befejeződik."
}
},
"refused_by_busy_guard": false,
"controller_says": "2026/09/22 19:45:20 delete.go:542: [ERROR] [stacks] RemoveStack privatebin REFUSED (busy): a backup or restore is running (single-flight held)"
}
RC=0
@@ -0,0 +1,39 @@
#!/usr/bin/env python3
"""B2 — prove the BUSY guard specifically (R-633).
B's first run refused the remove with `409 still running — stop it first`, which is the PRE-EXISTING
running check, not v0.262.0's new guard: the restore had already finished. To reach the new guard the
app must be STOPPED (so the running check passes) while the backup side is busy.
"""
import json, os, sys, time
HERE = os.path.dirname(os.path.abspath(__file__))
sys.path.insert(0, os.path.join(os.path.dirname(HERE), "update-night-2026-09-21"))
import walk as w # noqa: E402
OUT = json.load(open(os.path.join(HERE, "live262.json")))
app, sub = "privatebin", "paste"
rec = {"scenario": "B2 (R-633) — the BUSY guard, reached deliberately", "app": app}
w.login()
w.deploy(app, sub)
w.wait_app(sub, "/", tries=30)
# stop it, so the pre-existing "still running" check cannot answer first
w.ctl("POST", f"/api/stacks/{app}/stop")
for _ in range(24):
time.sleep(5)
if w.stack(app).get("state") != "running":
break
rec["state_before"] = w.stack(app).get("state")
# box-wide backup: `POST /api/backup/run` holds the single-flight the guard consults
code, d = w.ctl("POST", "/api/backup/run")
rec["backup_started"] = {"http": code, "answer": str(d)[:120]}
time.sleep(3)
code, d = w.ctl("POST", f"/api/stacks/{app}/remove", {"remove_hdd_data": False, "remove_backups": False})
rec["remove_during_backup"] = {"http": code, "answer": d}
rec["refused_by_busy_guard"] = (code == "409" and "ment" in json.dumps(d, ensure_ascii=False).lower())
rec["controller_says"] = w.guest(
f"docker logs --since 5m felhom-controller 2>&1 | grep -iE 'RemoveStack {app}.*(REFUSED|busy)' | tail -3").strip()
OUT["B2"] = rec
json.dump(OUT, open(os.path.join(HERE, "live262.json"), "w"), ensure_ascii=False, indent=2)
print(json.dumps(rec, ensure_ascii=False, indent=2))
@@ -0,0 +1,19 @@
21:51:24 [1] privatebin already deployed — reusing
{
"scenario": "B2 (R-633) — the BUSY guard, reached deliberately",
"app": "privatebin",
"state_before": "stopped",
"backup_started": {
"http": "200",
"answer": "{'ok': True, 'message': 'Mentés elindítva'}"
},
"remove_during_backup": {
"http": "409",
"answer": {
"ok": false,
"error": "Az alkalmazáson mentés vagy visszaállítás fut. Várd meg, amíg befejeződik."
}
},
"refused_by_busy_guard": true,
"controller_says": "2026/09/22 19:49:46 delete.go:557: [ERROR] [stacks] RemoveStack privatebin REFUSED (busy): a backup or restore is running (single-flight held)\n2026/09/22 19:51:32 delete.go:557: [ERROR] [stacks] RemoveStack privatebin REFUSED (busy): a backup or restore is running (single-flight held)"
}
@@ -0,0 +1,49 @@
21:41:22 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
21:41:32 [1] deployed, controller state=running, pinned={'privatebin': 'privatebin/pdo:2.0.6'}
21:41:32 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'}
21:41:57 [4] backup idle; last=None
21:41:57 [R] restoring privatebin from snapshot 'helyi' (of 1 offered)
21:41:57 [R] POST /backup/restore -> HTTP/2 302 ['location: /backups/restore?flash=flash.restore.started']
21:41:57 + 0.0s restore (True, None, None)
21:42:07 + 10.1s restore (False, None, None)
21:42:07 [R] after restore: state=running hold=None phase=None
== B
{
"scenario": "B (R-633/R-626)",
"app": "privatebin",
"snapshots": 1,
"restore_started": "{'ok': True, 'snapshot_id': 'helyi', 'snapshots': [{'time': '2026-09-22T19:41:55Z', 'short_id': 'helyi', 'tier': 1, 'drive_label': 'Belső SSD (rendszer)'}], 'http': 'HTTP/2 302', 'location': ['locatio",
"remove_during_restore": {
"http": "409",
"answer": {
"ok": false,
"error": "stack \"privatebin\" is still running — stop it first before removing"
}
},
"refused_as_expected": true,
"remove_after_restore": {
"http": "200",
"verified": true,
"reappeared": null
},
"containers_60s_later": ""
}
== C
{
"scenario": "C (R-634)",
"app": "sparkyfitness",
"setup": "total 28\ndrwxr-xr-x 2 root root 4096 Sep 22 19:43 .\ndrwxr-xr-x 57 root root 4096 Sep 13 20:22 ..\n-rw-r--r-- 1 root root 4737 Sep 20 14:32 .felhom.yml\n-rw-r--r-- 1 root root 41 Sep 22 19:43 app.yaml\n-rw-r--r-- 1 root root 4936 Sep 22 19:43 docker-compose.yml",
"before": {
"deployed": false,
"state": "not_deployed"
},
"remove": {
"http": "200",
"answer": "{'ok': True, 'data': {'removed': 'sparkyfitness', 'volumes_removed': [], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'backup_paths_removed': ['/mnt/sys_drive/felhom-data/backups/primary/sparkyfitness (48K)'], 'verified': True}, 'message': 'Stack sparkyfitness removed'}"
},
"accepted": true,
"leftovers": "NONE"
}
RC=0
@@ -0,0 +1,85 @@
#!/usr/bin/env python3
"""B (R-633) and C (R-634) live on 9202, against v0.262.0."""
import json, os, sys, time
HERE = os.path.dirname(os.path.abspath(__file__))
sys.path.insert(0, os.path.join(os.path.dirname(HERE), "update-night-2026-09-21"))
import walk as w # noqa: E402
OUT = json.load(open(os.path.join(HERE, "live262.json"))) if os.path.exists(
os.path.join(HERE, "live262.json")) else {}
def save(tag, rec):
OUT[tag] = rec
json.dump(OUT, open(os.path.join(HERE, "live262.json"), "w"), ensure_ascii=False, indent=2)
print("\n== " + tag + "\n" + json.dumps(rec, ensure_ascii=False, indent=2)[:1500], flush=True)
def B():
"""A remove sent DURING a restore must be refused; a remove after it must be VERIFIED."""
# gokapi is crash-looping from this morning's R-633 artefact — its config volume was removed, so
# its binary can never start and the stack never settles. A subject that cannot reach a steady
# state proves nothing about a guard that fires between steady states. privatebin is light,
# reliable and was walked all night.
app, sub = "privatebin", "paste"
rec = {"scenario": "B (R-633/R-626)", "app": app}
w.deploy(app, sub)
w.wait_app(sub, "/", tries=40)
w.backup_now(app)
snaps = w.snapshots(app)
rec["snapshots"] = len(snaps)
# start a restore and, while it runs, ask to remove
r = w.restore(app)
rec["restore_started"] = str(r)[:200]
time.sleep(2)
code, d = w.ctl("POST", f"/api/stacks/{app}/remove", {"remove_hdd_data": False, "remove_backups": False})
rec["remove_during_restore"] = {"http": code, "answer": d}
rec["refused_as_expected"] = code not in ("200", "202")
# let the restore settle, then remove properly
for _ in range(60):
st = w.stack(app)
if st.get("state") in ("running", "unhealthy", "stopped", "not_deployed"):
break
time.sleep(5)
time.sleep(10)
code, d = w.ctl("POST", f"/api/stacks/{app}/stop")
time.sleep(8)
code, d = w.ctl("POST", f"/api/stacks/{app}/remove", {"remove_hdd_data": True, "remove_backups": True})
rec["remove_after_restore"] = {"http": code, "verified": (d.get("data") or {}).get("verified"),
"reappeared": (d.get("data") or {}).get("reappeared_removed")}
time.sleep(30)
rec["containers_60s_later"] = w.guest(
f"docker ps -a --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'").strip()
save("B", rec)
def C():
"""A half-state — app.yaml on disk, deployed=false, no containers — must be REMOVABLE."""
app = "sparkyfitness"
rec = {"scenario": "C (R-634)", "app": app}
# build the exact measured shape by hand on the scratch guest: the stack dir with an app.yaml
# and no containers, which is what two failed deploys left behind twice.
rec["setup"] = w.guest(f'''set -e
mkdir -p /opt/docker/stacks/{app}
cp /var/lib/docker/volumes/felhom-controller-data/_data/data/catalog-cache/templates/{app}/docker-compose.yml /opt/docker/stacks/{app}/ 2>/dev/null || true
printf 'desired_state: stopped\\ninstalled_images:\\n' > /opt/docker/stacks/{app}/app.yaml
ls -la /opt/docker/stacks/{app}/''').strip()[-300:]
w.ctl("POST", "/api/stacks/rescan")
time.sleep(6)
st = w.stack(app)
rec["before"] = {"deployed": st.get("deployed"), "state": st.get("state")}
code, d = w.ctl("POST", f"/api/stacks/{app}/remove", {"remove_hdd_data": False, "remove_backups": True})
rec["remove"] = {"http": code, "answer": str(d)[:300]}
rec["accepted"] = code == "200"
time.sleep(5)
rec["leftovers"] = w.guest(f"ls /opt/docker/stacks/{app}/app.yaml 2>/dev/null || echo NONE").strip()
save("C", rec)
if __name__ == "__main__":
w.login()
for fn in (B, C):
try:
fn()
except Exception as e:
save(fn.__name__ + "-ERROR", {"error": f"{type(e).__name__}: {e}"})
File diff suppressed because one or more lines are too long