v0.262.0 + v0.262.1 live on 9202: four scenarios proven, R-630/R-633/R-621/R-614 closed
gates / gates (push) Successful in 30s
gates / gates (push) Successful in 30s
A (R-630, the P1): paperless-ngx, same app same button - failed at +313.0s with the app stopped under v0.261.0, done at +53.4s now, with no "no probe container" warning because the explicit healthcheck.container resolved the target. B (R-633): the busy guard fires - RemoveStack REFUSED (busy): a backup or restore is running - and the live proof caught it answering HTTP 500, because router.go maps remove errors by grepping the error TEXT. v0.262.1 makes it a typed error and a 409; re-proven live. C (R-634 half): a half-state with app.yaml on disk and deployed=false answered 200, leftovers NONE. Under v0.261.0 the same call said "not deployed". F (R-614): phase done before the remove, no phase at all after redeploying the same name. Also: the first B run proved NOTHING and nearly went down as a pass - the refusal came from the pre-existing "still running" check, not the new guard. Recorded. 09 6.1 and 8.8, the capability map, and STATUS updated. R-625 and R-634's mechanism are named as owed, not half-done. Register 325. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
File diff suppressed because one or more lines are too long
@@ -653,7 +653,7 @@ closed by construction: nothing reports an update complete on the compose exit c
|
||||
| 4 | `pinning` | the previous definition is copied aside and journaled, then the pin advances | pin put back |
|
||||
| 5 | `pulling` | `compose pull` | **pin and definition PUT BACK** — nothing ran (Scenario E) |
|
||||
| 6 | `starting` | `compose up -d --remove-orphans` | stop + HOLD |
|
||||
| 7 | `verifying` | the `.felhom.yml` health check through the existing probe, or 60 s of every container running and none restarting; bounded by `update.health_timeout` | **stop + HOLD; the pin STAYS** — the migration may have run (Scenario F) |
|
||||
| 7 | `verifying` | the `.felhom.yml` health check through the existing probe **when it resolves to a container**, else 60 s of every container running and none restarting; bounded by `update.health_timeout` | **stop + HOLD; the pin STAYS** — the migration may have run (Scenario F) |
|
||||
| 8 | `done` | installed images recorded, journal cleared | — |
|
||||
|
||||
**The two knobs** (`controller.yaml`, operator-owned): `update.backup_max_age` (default `24h`) and
|
||||
@@ -879,6 +879,18 @@ headlessly (R-460).
|
||||
| E — the automatic night | **COMPLETE.** The unattended HOLD was produced at last (312.9 s), with no retry across two further passes. It needed the image store of §6.5 |
|
||||
| F — the remaining apps | **DONE 2026-09-22 (`audits/DRILL-the-28-2026-09-22.md`): the 28 apps no drill had ever touched were walked in ONE night, so the catalog is now **53 of 53 attempted**, not 25.** 26 of the 28 deployed; 6 proven; 5 inconclusive; 14 had no within-a-major edge upstream that night; 1 failed honestly (`outline 1.9.1 -> 1.10.1`, HELD with the right sentence); 2 could not be deployed, one of them (`plant-it`) **by design** — it is `lifecycle: abandoned` and the product's lifecycle gate refused it, proven live for the first time. **The cost is now known and it is not machine time:** three concurrent walks did 28 apps in about four hours, and the binding constraints were FIXTURES (only 6 of 28 had a non-browser route that both seeded and read back) and the fact that `POST /api/backup/run` is BOX-WIDE, so concurrent walks serialise on it. **This leg also added the night's biggest finding**, which no count would have produced: R-630 |
|
||||
|
||||
**The words "when it resolves to a container" are v0.262.0's, and they are the whole of R-630.**
|
||||
Before it, the probe branch had no exit: `findProbeContainer` returning `""` set a message and
|
||||
looped, while the settle path that judges an app declaring NO check sat in the outer `else`,
|
||||
unreachable. So a stack whose probe resolved to nothing could only ever time out — and
|
||||
`failAndHold` then stopped an app whose containers were all healthy. **A stack with no probe is not
|
||||
healthy and not failing; it is settled on container state (§3), and never a reason to stop a running
|
||||
app.** Proven live on paperless-ngx 2026-09-22: the identical Update that ended **`failed` at
|
||||
+313.0 s with the app stopped** now ends **`done` at +53.4 s**
|
||||
(`audits/v0262-live-2026-09-22/live262.json`). Which container is probed is now decidable too —
|
||||
exact stack name, then `healthcheck.container`, then a UNIQUE prefix, else nothing with the
|
||||
candidates logged; the old rule took the FIRST prefix match.
|
||||
|
||||
**What the night ADDED to this table, which none of the legs anticipated:** the `verifying` phase
|
||||
trusts the `.felhom.yml` probe absolutely, and three of 53 templates name a probe the app does not
|
||||
answer — so a SUCCESSFUL update of those apps ends by STOPPING a working app (**R-618**, P1). That is
|
||||
@@ -1121,6 +1133,16 @@ Version strings stay in the logs, the API and the hub.
|
||||
**So limitation 8 now has two shapes, not one:** a probe that names the wrong target (R-618,
|
||||
fixed) and **no probe at all** (**R-630, raised to P1**), and the static gate can see the first
|
||||
but not the second, because there is nothing to compare.
|
||||
**FIXED 2026-09-22 in controller v0.262.0**, and the fix is in the phase table above: the
|
||||
no-probe case now settles on container state instead of looping, and the probe TARGET is
|
||||
decidable (exact name → `healthcheck.container` → a UNIQUE prefix → nothing, candidates logged).
|
||||
`paperless-ngx` and `immich` carry the explicit field, and the catalog gate REFUSES a probe that
|
||||
resolves to nothing rather than warning about it. **So limitation 8's two shapes are both closed
|
||||
in the product**; what remains is that six templates still cannot be judged STATICALLY (R-631,
|
||||
all five read live and correct) and that a probe can still be right about the port and wrong
|
||||
about what a 200 means — `romm` answered 200 from nginx while its workers were being OOM-killed
|
||||
for six hours (**R-635**), which is a third shape again and the reason "the update is guarded"
|
||||
must never be read as "the new version runs".
|
||||
**Two more things the same night measured, both about state rather than health:** a `remove` sent
|
||||
while a restore is still running reports success and leaves a container restarting with a live
|
||||
public route (**R-633**) — and the product already has exactly that guard for `update` and for
|
||||
|
||||
Reference in New Issue
Block a user