THE TWENTY-EIGHT: every app no drill had touched, walked in one night
gates / gates (push) Successful in 27s
gates / gates (push) Successful in 27s
All 28 walked on scratch guest 9202 against the private drill catalog. 26 deployed, 6 proven, 5 inconclusive, 14 with no upstream edge, 1 failed honestly (outline 1.9.1->1.10.1, HELD with the right sentence), 2 undeployable - one (plant-it) by design, refused by the lifecycle gate, proven live for the first time. Each app also got the half the update night skipped: a restore from its own copy with the seed read back again - 21 restored, 2 correctly REFUSED per 07 6.2. R-630 RAISED TO P1 by measurement: a stack with NO probe container does not skip verifying - it waits out the full health timeout and HOLDS, stopping an app whose three containers read healthy. The controller's own words: "not healthy within 5m0s (last: no probe container)". R-633 opened: a remove sent during a restore reports success and leaves a container restarting with a live public route. The product already refuses that clash for update and for restore, naming the blocker; remove has no such guard. R-634 opened: an app can be running, healthy and serving while recorded as deployed=false, and is then unremovable. Reproducible alone on sparkyfitness; concurrency-linked on two others. R-631 and R-632 CLOSED. Register 321 -> 323. Seven interventions, six of them my own harness - named, with what each cost. No product code. The live catalog's image: lines are byte-identical to the start of the night. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -799,9 +799,11 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-628** | **[P2-MEDIUM] An empty search of a mailbox I do not control was turned into a claim about what a THIRD PARTY had done, and it went into the register as fact.** FOUND 2026-09-22; the operator caught it within minutes by producing the thread. The `due_checks_gate` fired R-433 (*have Hetzner answered?*). Two searches of the felhom catch-all came back empty, and the emptiness was written into R-433 as **"Hetzner has not answered"** and **"there is no evidence the tickets were ever opened"**. Both were false: ticket **#2026090103040671** had been opened and answered. **THE PRECISE FAULT IS THE INFERENCE, NOT THE QUERY — and that distinction is the whole value of this row.** Re-run afterwards WITH a positive control: `from:monitoring@felhom.eu` returns **201 threads**, `in:anywhere … includeTrash` reaches SENT, TRASH and mail back to January — **the instrument works.** And the exact ticket number, the exact subject `Storage Box issue`, and `from:hetzner` each still return **nothing**. So the literal finding — *this correspondence is not in this mailbox* — was CORRECT. What was invented was the step from there to *Hetzner has not answered* and *the tickets were never opened*. **A mailbox I can read is not the only place a reply can be**, and the operator's own account is exactly where a support ticket he opened would land. An absent record in ONE place can never answer a question about what SOMEONE ELSE did. **This is R-96 rule 3 in a new surface, one step further out than R-607**: there the instrument reported a stale value as current; here a correct observation was promoted to a conclusion it could not carry. **It is sharper still because the same session, the night before, gave every fixture a negative control and every gate a red-proof — and then reached for a search with neither.** **THE RULE, in one line:** *an empty search may be reported as "absent from the place I looked", never as "it did not happen" — and only after a control query that MUST hit has been seen to hit.* **Done:** the rule is written into the Gmail-access memory, where the next session meets it before it searches rather than after. | **CLOSED 2026-09-22 — rule recorded; R-433 corrected with the real answers** |
|
||||
| **R-627** | **[P2-MEDIUM] Nothing checked that the register is a well-formed table, so an append that ate two rows' state cells went unnoticed until a person read the file — and one row had been broken the same way for 45 days.** FOUND 2026-09-22. The 2026-09-21 update night appended measured results to eight rows with a regex that matched each row's trailing state cell; on **R-446** and **R-458** it consumed the cell and did not restore it, the cell reappeared as a stray FOURTH cell on a DUPLICATED copy of **R-626** and **R-625**, and a blank line was left between each pair. The register then reported **317 rows for 315 findings**, two rows carried no state at all, and two findings existed twice with contradictory state cells. **Nothing caught it:** `one_register_gate.py` compares this file against ROADMAP and `closed_register_gate.py` forbids an id in BOTH files — neither asks whether the file is a well-formed table, and neither notices an id duplicated WITHIN it. **THE RED-PROOF THEN FOUND AN OLDER INSTANCE NOBODY HAD SEEN: R-254 lost its state cell on 2026-08-08 (commit `59527d0`) and had rendered without a State column for 45 days.** **Closed the same day:** `scripts/register_shape_gate.py`, registered in `repo_gates.py` as gate 14 and reached by the pre-push hook, refusing a row that does not end with `|` (an eaten state cell), a duplicated id, or a blank line splitting the table; four decoys in `test_gate_decoys.py` — three convicting on the exact damage shapes and one asserting a healthy register still passes. **TWO THINGS THE RED-PROOF CORRECTED IN THE GATE ITSELF, kept because they are the finding's real content:** a first draft counted CELLS and convicted **125 innocent rows** — register cells carry literal `|` inside prose and shell snippets (`owner: CC | …`), so a row cannot be split on `|`, and a count that cannot be computed is not a check; and it skipped malformed rows before counting ids, reporting **5 duplicates where there were 2**. **Repaired:** both state cells restored from the stray cells that carried them, the two duplicate rows deleted, R-254's verdict sentence given its cell back, and **15 blank lines that split the register into 12 separate markdown tables** removed — every row's text byte-identical afterwards, proven by diff. **And the gate immediately earned itself:** the very next row inserted in this session (R-628) left a blank line behind and the gate refused it. | **CLOSED 2026-09-22 — gate 14, four decoys, register repaired 317→315** |
|
||||
| **R-629** | **[P2-MEDIUM] The drill catalog sent the operator 47 CI-failure alarms in one night, and the drill method that created it did not mention CI at all.** FOUND 2026-09-22 while checking a different mailbox question — which is the only reason it was found. `admin/app-catalog-drill` was created on 2026-09-21 by `POST /api/v1/repos/migrate` from the live catalog, and a migrated repo inherits `has_actions: true`. Every drill push therefore ran the catalog's CI workflow, which failed immediately (the drill repo carries the workflow but the run has no meaningful gate context), and **each failure mailed `admin@felhom.eu`**: *"[felhom CI] gates FAILED in admin/app-catalog-drill"*. **47 runs, 47 alarms, all overnight, all unread and flagged IMPORTANT.** **WHY THIS MATTERS MORE THAN THE NOISE:** that mailbox is the operator's alarm channel, and R-168 made CI mail the thing that notices a bypassed gate. A night of throwaway failures from a repo nobody must ever act on trains the reader to skim exactly the sender that must never be skimmed — and it did it on the night the same mailbox was also carrying real `offsite_snapshots_dropped` and `offsite_proof_empty` alarms. **The drill method wrote down the fences it needed** — private repo, no customer box may follow it, reset at teardown — **and said nothing about CI, because nobody had run a drill repo through a CI-enabled Gitea before.** **FIXED 2026-09-22:** `has_actions` set to **false** on the drill repo (verified by re-reading the repo: `has_actions: False, private: True`), and `09` §6.5 now carries it as a step of creating a drill repo rather than as a thing to notice afterwards. **What is NOT done:** the 47 mails are still in the operator's inbox, unread — deleting another person's mail is not mine to do, and they are named here so they can be cleared in one search: `subject:"gates FAILED in admin/app-catalog-drill"`. | **CLOSED 2026-09-22 — actions disabled, method updated; the 47 mails are the operator's to clear** |
|
||||
| **R-630** | **[P2-MEDIUM] `paperless-ngx`'s health probe has never run, on any box, and the badge can never go red.** FOUND 2026-09-22 by the new `probe-matches-compose` gate, which reported it as a WARNING while looking for something else. `findProbeContainer` (`controller/internal/stacks/healthprobe.go:297`) takes the container whose name EQUALS the stack name, else the first whose name has it as a PREFIX; paperless-ngx's containers are `paperless-webserver`, `paperless-postgres` and `paperless-redis`, and **none of them begins with `paperless-ngx`**. The function returns `""`, `:62` counts the stack in `skippedNoContainer` and `continue`s, and no probe is ever built for it. **WHY THIS IS WORSE THAN A WRONG PROBE, WHICH IS WHAT R-618 WAS:** a wrong probe is a FALSE RED — loud, visible, and it stopped an app, which is how it was found within one night. This is a SILENT ABSENCE. The app page shows whatever the container state alone says, nothing contradicts it, and **R-96 rule 3 is the exact shape: an absent alarm is equally consistent with `healthy` and with `never checked`.** **NOT MEASURED, and said so rather than assumed:** what the guarded update's `verifying` phase does for a stack with no probe target — whether it passes immediately or waits out `update.health_timeout` — has not been run. `paperless-ngx` IS installed on demo-hp, so it is answerable on a real box. **Two candidate fixes, neither taken here because both are product code and this session was forbidden it:** give the template a `container_name: paperless-ngx` on its webserver service (catalog-only, one line, but it renames a running container on every existing box); or let the probe fall back to the service the compose declares first, and SAY SO in the health detail. **What is safe to say today:** the gate names it on every push, so it cannot go back to being invisible. | **OPEN — P2; owner: CC; needs the `verifying` behaviour measured on demo-hp before a fix is chosen** |
|
||||
| **R-631** | **[P3-LOW] Five templates cannot be judged by the probe gate at all, and one is correct only by accident.** FOUND 2026-09-22 when `check-probe-matches-compose.py` was run over all 53. The gate's oracle is the probed service's own compose `healthcheck.test`; where that dials no loopback URL the gate has nothing to compare and reports a WARNING rather than a pass. **`crafty-controller`, `mealie` and `uptime-kuma`** run their healthcheck through a python or script helper (`ssl._create_unverified_context()`, `socket.create_connection`, `extra/healthcheck`), so the port is inside code the gate does not execute. **`vikunja`** has no compose healthcheck on its probed service at all. **`home-assistant`** is the interesting one: its probe path `/api/` differs from the compose's `/manifest.json`, and it is healthy today **only because its check is `type: api` with no `expect` block, which `probeHTTP` treats as "any response is healthy"** — add `expect: {status: 200}` to that template, a change that looks like a tightening, and the app goes permanently unhealthy and every successful update of it starts stopping it. **The gate warns on it for exactly that reason and refuses to call it a pass.** **Needs:** a live probe reading for each of the five on a scratch guest — deploy, read `GET /api/stacks/<n>`, compare with the front door — which is one rotation night's work and closes the last gap R-618 left. | **OPEN — P3; owner: CC; five apps, one live reading each** |
|
||||
| **R-632** | **[P3-LOW] Twenty-eight of the 53 templates have never been deployed by any update drill, so nothing is known about whether their updates work.** COUNTED 2026-09-22 against the 2026-09-21 sweep, which is the widest one ever run. **20 apps have a verdict record** (14 proven, 3 failed, 3 inconclusive, plus tandoor re-walked to proven on 2026-09-22); **4 more were deployed as props in the bad-days legs with no edge walked** (`bentopdf`, `glance`, `uptime-kuma`, `wishlist`); **1 was deployed only to measure its probe** (`wger`); and **28 have never been deployed at all**: `calcom`, `calibre-web`, `claper`, `code-server`, `crafty-controller`, `emby`, `ghost`, `gokapi`, `gramps-web`, `homebox`, `homepage`, `immich`, `jellyfin`, `kimai`, `komga`, `onlyoffice`, `outline`, `paperless-ngx`, `plant-it`, `plex`, `radarr`, `rallly`, `recipe-importer`, `seerr`, `sonarr`, `sparkyfitness`, `termix`, `wanderer`. **THIS IS NOT A COMPLAINT ABOUT THE SWEEP** — it took one app-catalog-wide night to go from 3 apps ever measured to 21, and a night is the unit available. It is a record of what the catalog's update promise currently rests on: **for 28 of 53 apps, nothing.** **The list is the nightly rotation's queue**, smallest and least stateful first; `paperless-ngx` should be early because R-630 needs a live reading from it anyway, and `crafty-controller`, `mealie` and `uptime-kuma` should be early because R-631 needs one from each. Machine-readable copy: `audits/probe-fix-2026-09-22/not-judged.json`. | **OPEN — P3; owner: CC; the rotation works this list, 28 apps** |
|
||||
| **R-630** | **[P2-MEDIUM] `paperless-ngx`'s health probe has never run, on any box, and the badge can never go red.** FOUND 2026-09-22 by the new `probe-matches-compose` gate, which reported it as a WARNING while looking for something else. `findProbeContainer` (`controller/internal/stacks/healthprobe.go:297`) takes the container whose name EQUALS the stack name, else the first whose name has it as a PREFIX; paperless-ngx's containers are `paperless-webserver`, `paperless-postgres` and `paperless-redis`, and **none of them begins with `paperless-ngx`**. The function returns `""`, `:62` counts the stack in `skippedNoContainer` and `continue`s, and no probe is ever built for it. **WHY THIS IS WORSE THAN A WRONG PROBE, WHICH IS WHAT R-618 WAS:** a wrong probe is a FALSE RED — loud, visible, and it stopped an app, which is how it was found within one night. This is a SILENT ABSENCE. The app page shows whatever the container state alone says, nothing contradicts it, and **R-96 rule 3 is the exact shape: an absent alarm is equally consistent with `healthy` and with `never checked`.** **NOT MEASURED, and said so rather than assumed:** what the guarded update's `verifying` phase does for a stack with no probe target — whether it passes immediately or waits out `update.health_timeout` — has not been run. `paperless-ngx` IS installed on demo-hp, so it is answerable on a real box. **Two candidate fixes, neither taken here because both are product code and this session was forbidden it:** give the template a `container_name: paperless-ngx` on its webserver service (catalog-only, one line, but it renames a running container on every existing box); or let the probe fall back to the service the compose declares first, and SAY SO in the health detail. **What is safe to say today:** the gate names it on every push, so it cannot go back to being invisible. **MEASURED 2026-09-22 (the twenty-eight), and the answer is the WORST of the three possibilities, so this row is RAISED P2 → P1.** paperless-ngx was deployed on guest 9202 with all three containers **`healthy`**, the controller reading **`running`** and the front door answering **302**. The Update was then pressed (no upstream edge exists tonight, so it was pressed on the same version — which is what a household does on an up-to-date app and still walks the whole phase machine; stated rather than glossed). Phases: `checking` → `safety-dump` → `pinning` → `pulling` → `starting` (+2.1 s) → `verifying` (+3.1 s) → **`failed` at +313.0 s, the app STOPPED** (`state_after: stopped`, front door **404**). **THE CONTROLLER'S OWN WORDS NAME THE CAUSE, and no inference was needed:** *`update paperless-ngx FAILED after the new version was started: not healthy: not healthy within 5m0s (last: no probe container) — stopping and HOLDING the app; the pin stays on the new version (its migration may have run)`*. **`no probe container`.** So `verifying` does not pass when there is no probe and it does not skip — **it waits out the full `update.health_timeout` and then HOLDS.** **WHAT THIS COSTS:** every paperless-ngx household that presses Update has their working app **stopped for five minutes and then left stopped**, and is sent to a restore they do not need — and the hold sentence correctly warns that the copy *„csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem"*, so for this class-A app the route back is the off-site copy. **This is R-618's outcome reached by the opposite road:** there a probe named the wrong port; here no probe exists at all, and the static gate cannot see it because there is nothing to compare. The gate does name it on every push as a WARNING, which is how it was found. **Needs (unchanged, and now urgent):** give the template a `container_name: paperless-ngx` on its webserver service, or make `verifying` treat "no probe target" as something other than a failure — both are product/catalog decisions. Evidence: `audits/the-28-2026-09-22/sidejobs/r630-controller-words.txt`, `sidejobs/r630.json`. | **OPEN — RAISED TO P1 2026-09-22 by measurement; owner: CC; a successful update STOPS the app** |
|
||||
| **R-631** | **[P3-LOW] Five templates cannot be judged by the probe gate at all, and one is correct only by accident.** FOUND 2026-09-22 when `check-probe-matches-compose.py` was run over all 53. The gate's oracle is the probed service's own compose `healthcheck.test`; where that dials no loopback URL the gate has nothing to compare and reports a WARNING rather than a pass. **`crafty-controller`, `mealie` and `uptime-kuma`** run their healthcheck through a python or script helper (`ssl._create_unverified_context()`, `socket.create_connection`, `extra/healthcheck`), so the port is inside code the gate does not execute. **`vikunja`** has no compose healthcheck on its probed service at all. **`home-assistant`** is the interesting one: its probe path `/api/` differs from the compose's `/manifest.json`, and it is healthy today **only because its check is `type: api` with no `expect` block, which `probeHTTP` treats as "any response is healthy"** — add `expect: {status: 200}` to that template, a change that looks like a tightening, and the app goes permanently unhealthy and every successful update of it starts stopping it. **The gate warns on it for exactly that reason and refuses to call it a pass.** **Needs:** a live probe reading for each of the five on a scratch guest — deploy, read `GET /api/stacks/<n>`, compare with the front door — which is one rotation night's work and closes the last gap R-618 left. **CLOSED 2026-09-22 — all five read live on guest 9202, and all five probes are CORRECT.** Each app was deployed, its listening sockets read from inside the probed container, and the probe's own target dialled **on the compose network**, which is the call the controller makes. `mealie` `tcp 9000` → listens `0.0.0.0:9000`, dial 200. `uptime-kuma` `http 3001` → listens `*:3001`, dial 302 (and `http` calls any response healthy, so 302 passes and proves something answers). `vikunja` `api 3456 /api/v1/info` **with `expect: {status: 200}`** → dial **200** — the one that could have failed, because its expect block compares the code. `crafty-controller` `tcp 8443` → the controller's own log reads `Health probe crafty-controller: TCP :8443 -> ok (1ms)` twice, six minutes apart. **`home-assistant` is the one to carry forward:** `api 8123 /api/` with NO expect → the dial returns **401**, not 200. It reads healthy only because `probeHTTP` treats any response as healthy for that shape (`healthprobe.go:253-262`). **Add `expect: {status: 200}` to that template — a change that looks like a tightening — and home-assistant goes permanently unhealthy and every successful update of it starts stopping it.** That is R-618 one edit away, now a measured number rather than a caution. **So the gate's WARN list is not a backlog of suspects: it is four correct templates the gate honestly cannot prove, plus one correct by accident.** Evidence: `audits/the-28-2026-09-22/sidejobs/r631.json`. | **CLOSED 2026-09-22 — five live readings, five correct probes; home-assistant's fragility is now a number** |
|
||||
| **R-632** | **[P3-LOW] Twenty-eight of the 53 templates have never been deployed by any update drill, so nothing is known about whether their updates work.** COUNTED 2026-09-22 against the 2026-09-21 sweep, which is the widest one ever run. **20 apps have a verdict record** (14 proven, 3 failed, 3 inconclusive, plus tandoor re-walked to proven on 2026-09-22); **4 more were deployed as props in the bad-days legs with no edge walked** (`bentopdf`, `glance`, `uptime-kuma`, `wishlist`); **1 was deployed only to measure its probe** (`wger`); and **28 have never been deployed at all**: `calcom`, `calibre-web`, `claper`, `code-server`, `crafty-controller`, `emby`, `ghost`, `gokapi`, `gramps-web`, `homebox`, `homepage`, `immich`, `jellyfin`, `kimai`, `komga`, `onlyoffice`, `outline`, `paperless-ngx`, `plant-it`, `plex`, `radarr`, `rallly`, `recipe-importer`, `seerr`, `sonarr`, `sparkyfitness`, `termix`, `wanderer`. **THIS IS NOT A COMPLAINT ABOUT THE SWEEP** — it took one app-catalog-wide night to go from 3 apps ever measured to 21, and a night is the unit available. It is a record of what the catalog's update promise currently rests on: **for 28 of 53 apps, nothing.** **The list is the nightly rotation's queue**, smallest and least stateful first; `paperless-ngx` should be early because R-630 needs a live reading from it anyway, and `crafty-controller`, `mealie` and `uptime-kuma` should be early because R-631 needs one from each. Machine-readable copy: `audits/probe-fix-2026-09-22/not-judged.json`. **WORKED IN ONE NIGHT, 2026-09-22 — all 28 walked, so this row CLOSES and hands its findings to others.** Every one was installed on guest 9202 against the private drill catalog and taken through the same walk: deploy at the live pin, seed through the app's own front door, read it back, „Mentés most”, the guarded Update where a real within-a-major edge exists upstream, **restore from that copy and read the seed back a second time** (the half the update night skipped), then remove and a 60-second check that nothing came back. **26 of 28 deployed; 6 proven; the rest inconclusive, no-edge or refused.** **What the night produced that this row could not have predicted:** R-630 raised to P1 by measurement (a stack with no probe container has its working app STOPPED by a successful update), R-633 (a remove during a restore leaves an orphan with a live public route), R-634 (an app running and healthy while recorded as not deployed, and then unremovable), and one real upstream edge that HELD honestly (`outline 1.9.1 → 1.10.1`). **Also settled:** `plant-it` is `lifecycle: abandoned` and the product refuses to install it — the only lifecycle-gated template in the catalog, and its gate is now proven live. Full record: `audits/DRILL-the-28-2026-09-22.md`. | **CLOSED 2026-09-22 — all 28 walked in one night; the findings live in R-630, R-633, R-634** |
|
||||
| **R-633** | **[P2-MEDIUM] A remove sent while a restore is still running reports success, deletes the app's record, and leaves a container restarting forever with a live public route.** MEASURED 2026-09-22 on guest 9202, controller v0.261.0, during the twenty-eight walk. `gokapi` was restored from its own local copy at 11:34:07 and removed at 11:34:22. `POST /backup/restore` answers **302 and works in the background**; the remove tore down what existed and the restore's own `compose up` then RE-CREATED the container at **11:34:24**. **Both calls returned success.** Twenty-five minutes later: `GET /api/stacks/gokapi` reads **`deployed: false`**, and `docker ps -a` shows `gokapi` **`Restarting (1)`** with `RestartCount` climbing, carrying its full traefik label set — including `traefik.http.routers.gokapi.rule: Host(`.enkisfelhom.hu`)`, **a rule with an empty subdomain**, because the deploy values that filled it were deleted with the app. Its own log loops *„Salt for admin password invalid, generating new salt… password does not appear to be a SHA-1 hash"* — the volume holding its config was removed correctly, so the binary can never start. **WHY THIS IS A ROW AND NOT A HARNESS ARTEFACT:** a household can press exactly these two buttons in exactly this order, the product accepted both, and **the remove reported success while leaving the orphan**. Presence of a success message is not evidence of a result. **WHAT IT COSTS:** an app the household believes is gone keeps a container in a restart loop, keeps a route registered on the public reverse proxy, and is invisible to every product surface because the controller no longer records the stack. Nothing in the alarm ladder fires: `08` §4 keys on stacks the controller KNOWS about. **This is R-626's class with the mechanism finally visible** — that row recorded a removed `navidrome` coming back and could not diagnose it because the controller had restarted; here the window is 17 seconds and both halves are in the evidence. **Needs:** the remove path to refuse, or to wait, while a restore for the same stack is in flight — and, either way, to verify the teardown rather than report success without looking. Evidence: `audits/the-28-2026-09-22/apps/gokapi/came-back-evidence.txt`, `apps/gokapi/log.txt`. | **OPEN — P2; owner: CC; product code, so not fixed in this unattended run** |
|
||||
| **R-634** | **[P1-HIGH] An app can be RUNNING, HEALTHY and serving while the controller records it as not deployed — and in that state the household cannot remove it through the product at all.** MEASURED 2026-09-22 on guest 9202, controller v0.261.0, on **two independent apps in one night**: `outline` and `sparkyfitness`. **The controller's own words, in order.** `outline` deploy accepted 11:54:47; **`11:55:51 StopStack outline: current state=deploying deployed=true containers=0`** — a stop while the stack is still deploying; `11:56:09 SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, **0 encrypted**, 3 sensitive fields` (the two saves before it both read `3 encrypted`); then **`11:57:16` and `12:02:16 Health probe outline: API GET :3000/_health -> 200`** — the app is up and answering its own health endpoint; and **`12:02:19 StopStack outline: current state=running deployed=false containers=3`**. Three containers, state `running`, health 200, **`deployed=false`**. The remove then answers **`RemoveStack outline: state=not_deployed, deployed=false, orphaned=false, deploying=false` -> `[ERROR] Remove failed for outline: stack "outline" is not deployed`** for BOTH the remove-with-data and the remove-keeping-data call, while `ScanStacks` goes on finding the stack every ten seconds. `app.yaml` survives with `desired_state: stopped` and an **empty `installed_images`**. `sparkyfitness` produced the identical shape 13 minutes earlier. **WHAT THIS COSTS A HOUSEHOLD:** an app that works is invisible to the product as an installation — no badge, no update, no backup selection, and **no way to delete it**; the only exit is a shell. It is the mirror of R-633 (there the record is gone and the container remains; here the container is fine and the record is gone) and it is the **worse** of the two, because the app is serving customer traffic the whole time. **THE MECHANISM IS NOT DIAGNOSED, and this row says so rather than guessing.** What was tried: the controller's full container log for both apps (the sequence above), `app.yaml` on disk, `docker ps -a`, and `GET /api/stacks/<n>`. What was NOT done: reading `runComposeDeploy`'s pin-write path — this was an unattended run and the brief forbade product code. **The one discriminator worth running first:** both apps were walked while two other walks ran concurrently, and `POST /api/backup/run` is box-wide, so a backup or restore for a NEIGHBOURING app was in flight. A serial re-walk is queued tonight; if it reproduces alone, concurrency is not the cause. **A THIRD THING THE SAME LOG SHOWS, recorded here because it is one line away:** at `11:55:57` the health probe dialled **`http://outline-postgres:3000/_health`** — during startup, with the exactly-named container not yet running, `findProbeContainer`'s PREFIX fallback latched onto the POSTGRES sidecar and probed port 3000 on it. Transient, and it resolved once `outline` came up, but it is the same function R-630 is about. Evidence: `audits/the-28-2026-09-22/apps/half-state-outline-sparkyfitness.txt`. **THE SERIAL RE-WALK WAS RUN THE SAME NIGHT AND IT SPLITS THIS ROW IN TWO — recorded here rather than left as the first reading.** Walked again one at a time, with no other walk running: **`outline` deployed normally and removed clean**, and **`crafty-controller` deployed normally, updated `4.10.7 → 4.11.0` to `done`, restored and removed clean.** So for those two the half-state did NOT reproduce alone, and concurrency — a box-wide `POST /api/backup/run` or a restore in flight for a NEIGHBOURING app — is implicated rather than the deploy path itself. **`sparkyfitness` reproduced EXACTLY, alone, in 534 s**: deploy accepted, never reached `deployed` with a pin, `app.yaml` left with `desired_state: stopped` and an empty `installed_images`, and **both remove calls refused with `stack "sparkyfitness" is not deployed`.** **So the row stands, at one reproducible app instead of three, and the honest split is:** (a) `sparkyfitness` has a deploy that does not finish and leaves a record the product cannot clear — reproducible, P1; (b) under concurrent work the same unremovable half-state can be reached by apps that are otherwise fine, which is the more alarming half because those apps were **running, healthy and serving** while recorded as not deployed. **Neither half is diagnosed** — `runComposeDeploy`'s pin write was not read, because the brief forbade product code. | **OPEN — P1; owner: CC; reproducible alone on `sparkyfitness`, concurrency-linked on the other two; next step is `runComposeDeploy`'s pin write** |
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
Clearing a row means the check was DONE and its result recorded in that R-row —
|
||||
|
||||
Reference in New Issue
Block a user