THE TWENTY-EIGHT: every app no drill had touched, walked in one night
gates / gates (push) Successful in 27s
gates / gates (push) Successful in 27s
All 28 walked on scratch guest 9202 against the private drill catalog. 26 deployed, 6 proven, 5 inconclusive, 14 with no upstream edge, 1 failed honestly (outline 1.9.1->1.10.1, HELD with the right sentence), 2 undeployable - one (plant-it) by design, refused by the lifecycle gate, proven live for the first time. Each app also got the half the update night skipped: a restore from its own copy with the seed read back again - 21 restored, 2 correctly REFUSED per 07 6.2. R-630 RAISED TO P1 by measurement: a stack with NO probe container does not skip verifying - it waits out the full health timeout and HOLDS, stopping an app whose three containers read healthy. The controller's own words: "not healthy within 5m0s (last: no probe container)". R-633 opened: a remove sent during a restore reports success and leaves a container restarting with a live public route. The product already refuses that clash for update and for restore, naming the blocker; remove has no such guard. R-634 opened: an app can be running, healthy and serving while recorded as deployed=false, and is then unremovable. Reproducible alone on sparkyfitness; concurrency-linked on two others. R-631 and R-632 CLOSED. Register 321 -> 323. Seven interventions, six of them my own harness - named, with what each cost. No product code. The live catalog's image: lines are byte-identical to the start of the night. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
File diff suppressed because one or more lines are too long
@@ -877,7 +877,7 @@ headlessly (R-460).
|
||||
| C — the PostgreSQL rehearsal | **DONE and COSTED**: ~9 s of engine work, 155.9 s end to end for 49 MB / 48 tables. `pg_upgrade` still owed and may prove unnecessary |
|
||||
| D — the downgrade refusal | already done, v0.260.0 |
|
||||
| E — the automatic night | **COMPLETE.** The unattended HOLD was produced at last (312.9 s), with no retry across two further passes. It needed the image store of §6.5 |
|
||||
| F — the remaining apps | ~34 still unwalked. The fixtures for 20 exist and amortise |
|
||||
| F — the remaining apps | **DONE 2026-09-22 (`audits/DRILL-the-28-2026-09-22.md`): the 28 apps no drill had ever touched were walked in ONE night, so the catalog is now **53 of 53 attempted**, not 25.** 26 of the 28 deployed; 6 proven; 5 inconclusive; 14 had no within-a-major edge upstream that night; 1 failed honestly (`outline 1.9.1 -> 1.10.1`, HELD with the right sentence); 2 could not be deployed, one of them (`plant-it`) **by design** — it is `lifecycle: abandoned` and the product's lifecycle gate refused it, proven live for the first time. **The cost is now known and it is not machine time:** three concurrent walks did 28 apps in about four hours, and the binding constraints were FIXTURES (only 6 of 28 had a non-browser route that both seeded and read back) and the fact that `POST /api/backup/run` is BOX-WIDE, so concurrent walks serialise on it. **This leg also added the night's biggest finding**, which no count would have produced: R-630 |
|
||||
|
||||
**What the night ADDED to this table, which none of the legs anticipated:** the `verifying` phase
|
||||
trusts the `.felhom.yml` probe absolutely, and three of 53 templates name a probe the app does not
|
||||
@@ -1110,6 +1110,23 @@ Version strings stay in the logs, the API and the hub.
|
||||
unknown for six.**
|
||||
**AND THE SWEEP'S REAL CEILING, counted rather than felt: 28 of the 53 templates have never been
|
||||
deployed by any drill** (**R-632**) — the widening above went from 3 apps to 21, and 21 is not 53.
|
||||
**CLOSED THE NEXT NIGHT, 2026-09-22: all 28 were walked** (`audits/DRILL-the-28-2026-09-22.md`),
|
||||
so every template in the catalog has now been attempted at least once. **And the walk that closed
|
||||
it found something the probe work had left open.** `paperless-ngx` has no container whose name
|
||||
matches its stack name, so `findProbeContainer` returns nothing and its probe has **never run on
|
||||
any box**. Asked what `verifying` does with no probe to wait on, the answer is the worst of the
|
||||
three: it waits out the full `update.health_timeout` and **HOLDS**, stopping an app whose three
|
||||
containers all read `healthy`. The controller names it itself — *`not healthy within 5m0s (last:
|
||||
no probe container) — stopping and HOLDING the app`* — at **+313.0 s**, front door 404 afterwards.
|
||||
**So limitation 8 now has two shapes, not one:** a probe that names the wrong target (R-618,
|
||||
fixed) and **no probe at all** (**R-630, raised to P1**), and the static gate can see the first
|
||||
but not the second, because there is nothing to compare.
|
||||
**Two more things the same night measured, both about state rather than health:** a `remove` sent
|
||||
while a restore is still running reports success and leaves a container restarting with a live
|
||||
public route (**R-633**) — and the product already has exactly that guard for `update` and for
|
||||
`restore`, which name the blocking operation, but not for `remove`; and an app can be **running,
|
||||
healthy and serving while recorded as `deployed: false`**, in which state the product refuses to
|
||||
remove it at all (**R-634**, reproducible alone on `sparkyfitness`).
|
||||
9. **The hub does not record image tags at all.** Its report's container payload carries name, state,
|
||||
CPU and memory, and no image field (spike §5). So the fleet view of §6 slice 7 needs a hub-side
|
||||
change; it is not derivable from what is already reported.
|
||||
|
||||
Reference in New Issue
Block a user