The drill is complete. Teardown done in three layers plus Gitea; the live catalog's every `image:` line is proven identical to before. WHAT WAS MEASURED. 21 edges across 19 apps, on scratch guest 9202 through the product's own guarded Update, against a PRIVATE DRILL CATALOG so the live catalog carried no test reference at any point: 14 proven, 3 failed, 4 inconclusive. Each app seeded and read back through its OWN front door, with a negative control on every readback. Ten of the fourteen printed a verbatim migration line. Up from the three apps this project had ever measured. THE RESULT THAT MATTERS. R-618, P1: three of the 53 templates name a health probe the app does not answer, and because the guarded update WAITS on that same probe, a SUCCESSFUL update ends by STOPPING a working app. tandoor was measured serving HTTP 200 on the new version at four samples across five minutes, docker's own healthcheck green, and was then stopped and the household sent to a restore they did not need. zipline and wger are the same defect, both confirmed live. The gate that catches all three is static and cheap: both health checks already sit in the same file. WHAT THE NIGHT ANSWERED that was open. The UNATTENDED HOLD (312.9 s, pressed once, never again) — which needed a purpose-built image store, because the rule that makes automatic updates safe is the same rule that refuses the obvious way to break one. MariaDB across a major through the real button, all four observables, first time. PostgreSQL across a major, refusing exactly as predicted, with the conversion costed at ~9 s of engine work. There is NO single-flight: five updates ran at once and all ended honest. And the two EARLY power-cut phases nobody had cut in. TWELVE NEW ROWS (R-615..R-626), register 303 -> 315, and eight existing rows updated with what was measured — including two CORRECTIONS: R-606 records the pre-flight refusals as reaching an English household in English and they do not, and R-446/R-458 are both narrower than their rows state. Two instrument fixes were needed before anything could be trusted: the unattended caller turned every success into a timeout (R-623), and one of my own reproductions was wrong and is kept labelled with what it actually measured. Interventions: zero. No controller, agent or hub code written. The hub was never touched beyond the floor the operator asked for. Gates: repo_gates.py --fast, all 15 OK. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
36 KiB
Rows minted by the update night, 2026-09-21 — staging, to be pasted into OPEN-ITEMS.md
Highest existing id at the start of the night: R-614 (304 rows).
| R-615 | [P3-LOW] Pointing a box at a different app catalog by git.repo_url alone is INERT — the box keeps fetching from the repository it first cloned. FOUND 2026-09-21 by reading sync.go before running it, which is the only reason the update night's drill catalog worked at all. Syncer.gitCloneOrPull (controller/internal/sync/sync.go:274-306) clones only when <data>/catalog-cache/.git is absent; on every later cycle it runs git fetch --depth 1 origin <branch> + git reset --hard origin/<branch> against the remote stored in the clone, which buildRepoURL wrote at clone time. Changing git.repo_url in controller.yaml and restarting therefore changes nothing: the sync keeps pulling the old catalog and reports success. Measured: after the repoint, git -C <data>/catalog-cache remote -v still read app-catalog-felhom.eu; the box only followed the drill repo once the cache directory was removed. Why it matters beyond a drill: this is the one knob that would move a box to a different or a staged catalog — for a migration, a per-customer catalog, or a rollback of the catalog itself — and it silently does not work. Nothing is wrong with the CACHING, which is right; what is missing is that a changed repo_url must invalidate the clone. Fix shape: on start, compare git.repo_url with the clone's origin and re-clone when they differ (or git remote set-url + a full fetch); log which happened. A test that changes repo_url under an existing cache and asserts the next sync reads the NEW repo — it fails today. Evidence: audits/update-night-2026-09-21/04-9202-config-pre.txt, 05-9202-follows-drill.txt. | READY — rank P3-LOW; owner: CC (controller) |
| R-616 | [P3-LOW] The catalog credentials are stored in PLAINTEXT in the box's catalog clone and are printed by an ordinary git remote -v. FOUND 2026-09-21 on guest 9202 while pointing it at a private drill catalog. Syncer.buildRepoURL injects username:token into the HTTPS URL, and git clone persists that URL as the clone's origin, so <data>/catalog-cache/.git/config holds the token in the clear and any diagnostic that prints the remote leaks it — which is what happened in this session's own transcript, and is the same shape as R-580 (curl -w '%{redirect_url}'). maskRepoURL exists and is used for the LOG lines, so the masking intent is already there; the stored remote is the half that was missed. INERT ON THE FLEET TODAY — the live catalog is public and git.token is empty on every real box — which is exactly why it should be fixed before it is not: the day the catalog goes private, every box carries a readable credential and every support session that runs git remote -v prints it. Fix shape: store the remote WITHOUT credentials and supply them per-fetch (a credential helper, http.extraHeader, or GIT_ASKPASS), and a test asserting the clone's stored origin contains no @. Operator action from tonight, unrelated to the fix: the Gitea admin token used for the drill repo was printed by that command and must be rotated. Evidence: audits/update-night-2026-09-21/05-9202-follows-drill.txt (redacted). | READY — rank P3-LOW; owner: CC (controller); one operator action (rotate the Gitea admin token) |
| R-617 | [P3-LOW] The Gitea API token this project uses for pushes cannot create a repository through the documented endpoint, but CAN through repos/migrate — so "the token cannot do it" was nearly recorded as a fact when the truth was "one endpoint refuses it". FOUND 2026-09-21 creating the drill catalog. Both ~/.git-credentials tokens carry write:misc,write:notification,write:package,write:issue,write:repository; POST /api/v1/user/repos requires write:user and answers 403, and POST /api/v1/admin/users/<u>/repos requires write:admin and answers 403 too. POST /api/v1/repos/migrate with the same token answered 201 and created the private repository. Why this is a row and not a note: a session that stopped at the first 403 would have recorded "CC cannot create a Gitea repository" — an unfalsifiable capability claim of exactly the shape the workspace's standing rule 2 forbids — and every later drill would have been designed around a limit that does not exist. What it needs: one line in the operations notes saying which endpoint to use, and (optional, operator) a token scoped for the job so the migrate route is not load-bearing. Evidence: audits/update-night-2026-09-21/03-drill-repo.txt. | READY — rank P3-LOW; owner: CC (docs) |
| R-618 | [P1-HIGH] THREE apps are presented to the household as UNHEALTHY while they are working perfectly — and because the guarded Update waits on that same probe, a SUCCESSFUL update ends by STOPPING the working app and sending the household to a restore they do not need. RANK RAISED FROM P2 TO P1 BY A LIVE MEASUREMENT taken the same night, and the escalation is the whole point: tandoor's Update 2.6.13 → 2.6.15 was pressed at 21:16:47 and entered verifying at 21:17:46. At 21:18:28 the NEW version was Up 25 seconds and answering HTTP 200 on /accounts/login/ through the household's own front door — while the controller, probing port 8080 where nothing listens, could not see it. verifying therefore cannot pass, the full update.health_timeout is spent, Manager.failAndHold runs compose down, and the app is STOPPED. Nothing is lost — the data is in the volumes and the restore works — but one wrong port number in a template converts every successful update of that app into an outage plus an unnecessary restore, for every household running it. Evidence: audits/update-night-2026-09-21/14-tandoor-serving-while-verifying.txt. MEASURED 2026-09-21 on guest 9202 (controller v0.261.0, catalog f5f6a152b513). Two shapes, one class: (a) tandoor — the WRONG PORT. .felhom.yml probes port: 8080; the container listens on 80 and nothing else (ss -ltn inside it), the compose's own traefik label routes to 80, its own docker healthcheck reads healthy, and /accounts/login/ answers 200 through the household's real front door. GET /api/stacks/tandoor nevertheless reads state: "unhealthy". (b) zipline — the WRONG PATH. .felhom.yml probes /api/health, which zipline 4.6.1 answers 404 Route GET:/api/health not found; the compose healthcheck in the very same file uses /api/healthcheck and is correct and green. /dashboard answers 200. The controller reads unhealthy. This is the MIRROR of R-613 — that is a probe that passes on a broken app (a false GREEN, which no alarm catches); this is a probe that fails on a working app (a false RED). IT DOES NOT ALARM, AND THAT SETS THE RANK: 08-alarm-ladder.md §4 puts unhealthy deliberately in the NOT-down set, so no dead-app event and no customer mail follows — the damage is what the household READS, plus anything that gates on state. IT ALREADY COST A MEASUREMENT TONIGHT: this drill's harness waited for state == "running" and hung for its full budget on tandoor, an app that was up the whole time. An instrument waiting for a wrong answer looks exactly like a slow app. THE GATE THIS WANTS IS CHEAP AND STATIC, AND THAT IS THE FINDING'S REAL VALUE. Both halves of the answer live in the same template: compare the .felhom.yml probe's port and path against the compose's own healthcheck: test: URL. A sweep of all 53 templates on that rule was run tonight and returns five disagreements: tandoor (PORT — CONFIRMED live), zipline (PATH — CONFIRMED live), wger (PORT, probe 80 vs compose 8000 — CONFIRMED live the same night), home-assistant (PATH, /api/ vs /manifest.json — NOT MEASURED), and adventurelog (a FALSE POSITIVE of the sweep's own regex — it reads running live). So the rule finds both real defects, with two candidates and one false positive out of 53 — a good enough signal for a fast gate, provided it reports candidates rather than convictions and a person or a runtime check resolves them. The earlier, cruder rule (probe port vs the traefik port) is strictly worse: it clears zipline and convicts adventurelog. AND A SECOND FIX SHAPE, ON THE CONTROLLER SIDE, WORTH CONSIDERING BESIDE THE CATALOG ONE: in both confirmed cases the container's OWN docker healthcheck was green the whole time. A verifying phase that is about to stop a working app could ask that too — if the compose declares a healthcheck and docker reports healthy, the app is alive whatever our probe thinks. That does not excuse a wrong probe, but it turns this failure direction from an outage into a wrong label. It is a design question, not a defect, and is raised here rather than decided. Needs: fix tandoor's port (80), zipline's path (/api/healthcheck) and wger's port (8000) — all three are now CONFIRMED live, none is a guess; add the static gate with a decoy each way (R-421) — a template whose probe agrees must not read as a disagreement, and vice versa. THE GATE'S RULE WAS THEN SHARPENED BY READING healthprobe.go RATHER THAN ASSUMING IT, and the sharpening REMOVED a false conviction. type: http treats any response as healthy (healthprobe.go:258-261), and type: api with no expect block does the same (:265-268); only type: api WITH expect.status cares about the path or the code. So a PATH difference is a candidate only for the third shape, while a PORT difference is a candidate for all of them. Under that rule the 53-template sweep returns four candidates — tandoor, zipline and wger (all three CONFIRMED live — probe type: http, port: 80; inside the container port 80 is refused and port 8000 ANSWERED; docker's own healthcheck green; front door 302; the box reads unhealthy), and adventurelog (a false positive: its compose lists two containers' ports and the probe targets the backend; measured running). home-assistant is correctly CLEARED by the sharpened rule — type: api, no expect, so its /api/ answering 401 without a token is healthy, and its edge was PROVEN on the box tonight. The crude rule convicted it; the rule read from the code does not. That is the gate to build: two of 53 convicted, one suspected, one false positive, and the false positive is resolvable by one live check. Evidence: audits/update-night-2026-09-21/10-probe-port-sweep.txt, 12-probe-vs-compose-healthcheck.txt and 13-probe-sweep-sharpened.txt. | READY — rank P2-MEDIUM; owner: CC (catalog) |
| R-619 | [P3-LOW] A type: password deploy field is MANDATORY however required reads, and the deploy-fields contract says the opposite — so any caller that trusts it is refused. MEASURED 2026-09-21 on guest 9202 while widening the update drill. GET /api/stacks/grafana/deploy-fields serves {"env_var":"GF_SECURITY_ADMIN_PASSWORD","type":"password","generate":"password:16","required":false}; a deploy carrying only the two required:true fields is refused 400 „a(z) „Admin jelszó" mező kitöltése kötelező — használja a Generálás gombot…". The BEHAVIOUR is right and is a decision, not a bug: deploy.go:305-312 refuses a password field with no caller value on purpose — "We never silently auto-generate — the user needs to know their password" — which is the opposite of the secret case one branch above, where a generated value the customer never sees is exactly correct. The defect is the CONTRACT. .felhom.yml declares required: false, the API serves that verbatim, and nothing on the wire distinguishes "optional because the box will generate it" (secret) from "optional in the template and mandatory in the code" (password). A person using the deploy page never meets this because the page renders a Generálás button; anything that is not that page does, which now includes this drill harness and would include 09 §6.2's unattended caller the day it deploys anything. Fix shape (smallest that keeps the decision): serve required: true for type: password in the deploy-fields response — one place, derived rather than stored, so templates need no edit — and a test asserting a password field always reaches the wire as required. Alternatively state it in the field's description, which is weaker because it is prose. Evidence: audits/update-night-2026-09-21/apps/grafana/log.txt (the refusal) and batchA.log. | READY — rank P3-LOW; owner: CC (controller) |
| R-620 | [P3-LOW] A disabled notifier drops every event with NO local trace, so a box whose hub configuration is absent or broken stops telling anyone anything and leaves nothing behind that says so. FOUND 2026-09-21 on guest 9202 while trying to score the update night's alarm truth table. hub.enabled: false there, and Notifier.Publish returns at notify/notifier.go:269 — before any log line — as do NotifyHealthChange (:359) and four more entry points. Startup says it once ([INFO] Notifier disabled (hub not configured)) and then every later event, of every severity up to critical, vanishes without a word. The measurable consequence tonight: the whole event-and-mail half of the drill was structurally unmeasurable on this venue, and the alarm truth table below covers only the app page, the dashboard and the box's own log. That is a cost this session paid and named; the next one would pay it again. The consequence on a real box is smaller but not zero: the fleet's boxes have the hub enabled, and total silence is already caught by the hub's dead-man's-switch (staleness from the LAST REPORT, proven in the 2026-07-22 power-outage audit). What is NOT caught is the in-between — a box that still reports but whose notifier was disabled by a bad config push would go on reporting healthy while dropping every alarm, and the only evidence would be a single INFO line at the last restart. Fix shape: one DEBUG (or WARN, once per event type) line on the disabled path naming the event that was dropped, so the absence is visible where it happens rather than inferable from a startup line. Cheap, and it converts an invisible failure into a greppable one — R-96 rule 3 in the place that produces it. Evidence: audits/update-night-2026-09-21/11-notifier-disabled.txt. | READY — rank P3-LOW; owner: CC (controller) |
| R-621 | [P2-MEDIUM] A held update DESTROYS the evidence of why it failed: failAndHold runs compose down, the failing containers are removed, and their output is gone before anyone — household, operator or the next session — can read it. MEASURED 2026-09-21 on guest 9202 on a REAL upstream edge: adventurelog v0.12.1 → v0.13.0. The new backend applied nine Django migrations successfully and then never listened; the update held after the full 5-minute health wait. Manager.failAndHold (stacks/update.go:723) calls updateCompose(dir, env, "down"), which removes the containers rather than stopping them, and nothing captures their logs first. Within seconds the box's own log recorded Logs result for adventurelog: 0 bytes returned (empty) and docker ps -a held nothing at all. What survives is the WHAT and not the WHY: the controller line update adventurelog FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state unhealthy) and the household's sentence, both of which say the app did not come up and neither of which says the migrations ran and the server then failed to bind. This is R-320 ("evidence off the machine before the teardown") as a PRODUCT behaviour rather than a session habit — the teardown here is the product's own, it is correct to perform (a half-started new version must not keep running), and it happens before anyone can look. Why it matters beyond a drill: the hold sentence sends the household to a restore, and after the restore the only remaining question is should I press Update again? — which nobody can answer, because the one artefact that would say so no longer exists. It also makes every future held update unreportable to an upstream project. Fix shape: capture compose logs --no-color --tail N into the stack directory (beside applied-compose.yml, which already travels with the stack) IMMEDIATELY before the down, and surface it on the app page's hold panel or at least through the existing /api/stacks/<n>/logs fallback. Bounded size, written once per hold. A test that holds an app and asserts the captured file is non-empty — it fails today. Evidence: audits/update-night-2026-09-21/apps/adventurelog/why-it-failed.txt, state-after-hold.txt. | READY — rank P2-MEDIUM; owner: CC (controller) |
| R-622 | [P2-MEDIUM] adventurelog v0.13.0 migrates the customer's database and then does not serve — the edge must NOT be promoted, and it is the first real-catalog candidate this project has measured as unsafe. MEASURED 2026-09-21 on guest 9202 through the product's own guarded Update. v0.12.1 → v0.13.0 (backend AND frontend together, PostGIS held constant). The backend applied nine migrations, every one ... OK — adventures.0072_trail_wanderer_author_fields through integrations.0009_alter_endurainintegration_auth_method, plus billing.0001_initial — and then the container's own healthcheck failed with URLError: [Errno 111] Connection refused on five consecutive checks. The app never bound its port. The update held honestly after the full 5-minute wait. THE PRODUCT DID EVERYTHING RIGHT AND THAT IS HALF THE FINDING: the precondition found a Tier-1 copy one minute old, the safety dump was written, the pin advanced BEFORE the pull, the health wait was not short-circuited, the app was stopped rather than left half-running, and the hold sentence named the tier, the date and what the copy holds — „saját meghajtó, 2026-09-21 20:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza." This is exactly the case 09 §4 exists for: the migration RAN, so there is no undo, only a restore — and the restore is the thing slice 4 made sure existed first. What it needs: adventurelog stays OFF the promotion list; the cause is not diagnosed here (R-621 is why); and before it is ever promoted the edge should be re-run on the harness with its ABORT step, since an app that migrates and then refuses is the shape most likely to refuse the old image too. Evidence: audits/update-night-2026-09-21/apps/adventurelog/. | READY — rank P2-MEDIUM; owner: CC (catalog); NOT a Felhom defect — an upstream edge that fails |
Lines to ADD to existing rows (not re-filed)
R-607 — append: Seen again 2026-09-21 (update night), a dozen times in one session, and for the first time with a USER-VISIBLE consequence rather than a measurement one. On mealie the bump was pushed, POST /api/sync AND POST /api/stacks/rescan were both run, and the badge still read the up-to-date one (HU „Naprakesz", EN "Up to date") — the catalog had not reached catalog_images yet. The guarded Update was then pressed and reported „Frissitve" after 2.1 seconds having moved nothing at all: pinned, installed, the live compose line and docker inspect all still read v3.20.1. That is honest given a stale cache — the pin is written from the catalog's current definition, which was still the old one — but what the household sees is a button that says it updated them and did not. A NUMBER, at last, which is what this row asks for: the night's harness was changed to poll catalog_images until the pushed reference appears and to report how long that took; those figures are each edge's badge_catchup_seconds, and here they are: 4.4 s, 4.4 s, 4.5 s, 4.5 s — and 29.0 s. The four fast ones are one sync+rescan round; the 29-second one (nextcloud, an engine-sidecar bump) needed additional sync+rescan rounds before catalog_images carried the pushed reference. So the window is not a fixed scan interval — it varies by roughly 7x between edges on the same box in the same hour, which is why a caller (or a household) cannot know when the badge is safe to read. Before tonight this row had no number at all; it now has five, and they disagree with each other, which is itself the most useful thing about them. Every drill-catalog bump of the night was followed by POST /api/sync answering „Sablonok naprakészek — nincs változás" while the box's cache HAD moved, with catalog_images staying stale until a separate POST /api/stacks/rescan. The night's harness therefore rescans unconditionally after every sync, which is a workaround and not a fix. The window was still never measured as a NUMBER — that is what the row asks for and what remains owed.
R-462 — append: The count moved on 2026-09-21 from 3 apps to . The update night walked real, within-a-major upstream edges on guest 9202 through the product's own guarded Update, each seeded and read back through the app's own front door: .
proven, failed, inconclusive. Box-side fixtures for apps now exist at audits/update-night-2026-09-21/fixtures.py, and four of them (actualbudget, navidrome, audiobookshelf, vikunja) are ported into app-catalog-felhom.eu/scripts/upgrade_fixtures.py with seven new EDGES (U1–U7) so the same edges can be run on the harness venue with their ABORT step, which the box deliberately does not offer. Owed: the harness RUNS for those edges (the code is in; the runs are not), and fixtures for the apps recorded inconclusive tonight.
R-463 — append: Measured 2026-09-21 (update night), both halves. . And the conversion rehearsal Q5 asks for was costed on a real seeded datadir: .
R-469 / R-459 — append: The MariaDB engine major was pressed through the real Update BUTTON for the first time on 2026-09-21 (it had only ever been run on the harness). .
R-446 — append: Measured on the box 2026-09-21 (update night), leg B8. .
R-458 — append: Measured 2026-09-21 (update night), leg B9. .
R-613 — append: the update night could not seed uptime-kuma for the same reason and left it out rather than faking it.
R-460 — append: bookstack's edge was walked again on 2026-09-21 and is again half-proven — the database half read back through php artisan, the file half untouched. The limitation is unchanged and is now measured on the box as well as on the harness.
R-442 (CLOSED) — append, as a confirmation rather than a reopening: the fail-closed half was exercised again 2026-09-21 on guest 9202, where /api/disks answers agent not configured. Three apps deployed with an HDD_PATH (navidrome, audiobookshelf, romm) were each REFUSED at „remove with data" — „A(z) …/userdata/ tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó…" — with the app kept, and each was then removed successfully with the data KEPT. So the guard refuses the destructive half and leaves the non-destructive half available, which is exactly the shape the row describes. No change to the row's status.
| R-623 | [P3-LOW] The unattended-update caller turned every SUCCESS into a timeout, and then refused to press that app again — the instrument, not the box. FOUND 2026-09-21 (update night) by reading unattended-caller.py before relying on it for the Q4 hold measurement. Its call() returns the API envelope — {"ok": true, "data": {…}} — and follow() read update_phase and updating off the envelope, where neither exists. Both were therefore always None; the end test not updating and phase in ("done","failed") could never fire; every followed update ran the full 900-second timeout and was recorded timeout, which the caller treats as terminal and adds to never_again. main() unwraps data for the stack LIST, which is exactly why the within-a-major half of that night worked and this half did not. The 2026-09-21 run did not catch it because the only pass that reached follow() was Scenario F, whose log was lost to a buffering tail — the run's own honestly-recorded instrumentation gap turns out to have hidden a second one underneath it. This is the R-607 class in the evidence layer rather than the product layer: an instrument that can report a success as a timeout is not a measurement, and worse, it is a measurement that says the box behaved badly when the box behaved well. FIXED in the same file 2026-09-21 (unwrap data, with the reason written into the docstring so the next reader does not re-derive it), and the fixed caller is what produced tonight's unattended-hold leg. What it does NOT invalidate: the G-b no-retry proof, which is entirely in the refusal path and never reached follow(). What it DOES qualify: any future reading of that night's Scenario F timing — the "51 s – 1 m 26 s" figures come from the ATTENDED scenarios 04/05/07, not from the caller. | CLOSED 2026-09-21 — fixed in audits/update-arc-gaps-2026-09-21/unattended-caller.py |
| R-624 | [P3-LOW] Three of the catalog's apps cannot be seeded by ANY headless route, and for two of them that is a deliberate security decision — so the upgrade harness has a permanent ceiling nobody has written down. FOUND 2026-09-21 while widening R-462 from 3 apps to . vaultwarden and zipline close self-registration ON PURPOSE — vaultwarden by SIGNUPS_ALLOWED=false (R-512, „a stranger who guesses vault. must not be able to register"), zipline by answering E1037: User registration is disabled — and neither ships a CLI that could make an account instead. So there is no route to a first account without the admin secret, and that is correct: the harness must not be the reason a customer-facing app accepts strangers. gitea is a different and fixable case: the template sets no INSTALL_LOCK, so a fresh instance sits in its web-installer state and gitea admin user create refuses (MustInstalled() [F] Unable to load config file for a installed Gitea instance); POSTing the installer form first would work and was simply not written tonight. Why this is a row rather than three notes: 09 §3 decision 6 says the upgrade test goes to all apps, and R-462 is costed as if every app is reachable given enough fixture work. It is not. There is a class — apps whose only account-creating route the catalog deliberately closes — for which the honest maximum is inconclusive unless the harness is given the app's admin secret at deploy time, which is a decision nobody has taken. Needs: the class named in R-462's scope so the remaining count is honest; a decision on whether the harness may hold an app's admin secret (it already holds the ones IT generates — see R-619); and, separately and cheaply, a gitea installer-form fixture. Evidence: audits/update-night-2026-09-21/apps/{vaultwarden,zipline,gitea}/verdict.json. | READY — rank P3-LOW; owner: CC (catalog harness) |
R-606 — append: CONFIRMED 2026-09-21 (update night) on the HOLD sentence, which this row's own text ranks highest — „a PROMISE ABOUT WHETHER THE CUSTOMER'S FILES COME BACK". Read off /apps/adventurelog?lang=en while the app was genuinely held after a real failed upstream edge. Everything around it is correctly English — the nav, „An installed app is not running: AdventureLog (stopped)", „Update available — today", „Move to another storage" — and the two sentences that matter are Hungarian: „A(z) adventurelog frissítése … nem sikerült, és az alkalmazás nem indult el az új verzióval." and „Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 20:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza." So an English household is told in English that their app is stopped, and in Hungarian which copy brings their data back, when it was taken and what is inside it. Positive and negative controls both quoted. Also settled by the same reading, and it belongs to 09 Q4: the held app is surfaced on EVERY authenticated page, not only its own — the banner carried BOTH held apps at once. Evidence: audits/update-night-2026-09-21/15-r606-hold-sentence-on-the-english-page.txt and apps/adventurelog/held-page-en.html.
R-606 — append a SECOND correction: the row records v0.260.0 as having made the pre-flight REFUSALS reach an English household in English. Measured 2026-09-21: it did not. Three refusals were requested with ?lang=en, with the Hungarian request as the control, and all three came back identical Hungarian: held (the one that names which copy holds what), not_deployed, and disk („Nincs elég szabad hely a frissítéshez: 1.4 GB szabad…"). The mechanism is not a regression — it is that the pipe was built and the sentences never entered it. Router.langFor DOES honour ?lang= (api/i18n_api.go:30-40) and errText calls it; but MsgUpdateDiskFmt, MsgUpdateNotDeployed, MsgUpdateBusy and their siblings (stacks/update.go L81-96) are finished Hungarian string constants raised with fmt.Sprintf, and errText correctly renders "its own text" for an error carrying no bundle message. Only the sentences BORN as keys — v0.260.0's downgrade, v0.261.0's self_updating — actually translate. A row that records something as fixed when it is not is worse than an open row, which is why this correction is here rather than in prose. Evidence: audits/update-night-2026-09-21/20-refusals-in-english.txt.
R-458 — append: MEASURED 2026-09-21 (update night), leg B9, and the row's risk is NARROWER than it states. A .felhom.yml-only change (a health check for a path only a newer version would serve) was pushed to a FROZEN bentopdf — installed v2.8.6, catalog ahead. §5.4's asymmetry is confirmed live: the new .felhom.yml reached the box while the compose image: line stayed v2.8.6. But no false alarm was produced: ten samples over two minutes all read state=running with the front door at 200. The reason is the probe's own semantics, not luck — healthprobe.go:258-261 treats any response as healthy for type: http, and the bogus path answers 404, which is a response. So this row's false-alarm risk exists only for type: api probes carrying an expect block, where the status is compared; for every type: http template and every type: api without expect, a newer version's path is invisible to the probe. The row's actual claim — the failure direction is a false alarm, never data loss — stands and is now measured. Evidence: audits/update-night-2026-09-21/21-B9-frozen-app-newer-felhomyml.md.
| R-625 | [P2-MEDIUM] A HELD app keeps inviting the household to update it, and the button then refuses — the exact inconsistency R-524 removed for the other case, still present for this one. MEASURED 2026-09-21 (update night, leg B6). glance was HELD by a genuine unattended failed update. The drill catalog then published a fixed newer version — a real forward route, the thing a household would hope for. Afterwards the app page read „Frissítés elérhető — ma" / „Update available — today", in both languages, with the Update button offered; pressing it answered 409 reason='held' and the hold sentence. The BEHAVIOUR is correct and is DESIGN, not a defect: 09 §6.1 says the hold is settings.RestoreHold and that a successful unit restore lifts an update hold (only that kind) — a newer catalog version does not, and should not, because nobody has checked that the new version can start on data the failed one may have touched. The DEFECT is that the page says otherwise. R-524 settled precisely this shape for the Ahead case — "(a) is free and offers a household a downgrade … (c) show „Naprakész" and refuse the button … Why (c): the direction was already settled" — and chose to make the badge and the button agree. The HELD case still has them disagreeing, in the more painful direction: the badge invites, the button refuses, and the refusal is the same long sentence the household has already read. What it needs: the badge for a held app should say what is true — that the app is held and the way back is the restore — and the Update button should not be offered while RestoreHold stands. One verdict, read by both surfaces, exactly as R-524 did it. AND A QUESTION FOR 09 §3b Q4 THAT THIS MEASUREMENT RAISES AND DOES NOT ANSWER: the household's ONLY route out is a restore, even when the catalog has already shipped a fix. That is defensible, but it is now measured rather than assumed, and Q4's "does the box try again?" should be read next to it. Evidence: audits/update-night-2026-09-21/bad-days/B6-way-out-forwards/result.json. | READY — rank P2-MEDIUM; owner: CC (controller) |
R-446 — append: MEASURED ON A BOX 2026-09-21 (update night, leg B8), and it REFINES the row in two ways rather than merely confirming it. §8.1's numbers came from a registry sweep on DooPlex; this is the same question asked of a customer-shaped box, where the badge actually renders. On guest 9202, docmost's two floating pins were read as installed_images records them and compared against the upstream digests measured the same night: postgres:16-alpine → sha256:721873c34ceb9… on the box and sha256:721873c34ceb9… upstream, and redis:7-alpine → sha256:858f009f9709c… both sides. Identical. So the badge „Naprakész" is TRUE for this box, and the app reads correctly. (1) The defect's size is set by INSTALL AGE, not by the catalog. A floating pin is wrong only for a box that pulled BEFORE the tag moved; a box deployed after the repush holds the current image and its badge is right. R-446's "six repushed pins" measured the tag against the date the CATALOG set it, which is the right measure for the catalog and not for a box. (2) The producer Q6 needs ALREADY EXISTS on the box. installed_images records a real digest per service (installed.go §7.1) — the box knows exactly what it is running. What it cannot do is COMPARE, because the catalog carries no digest to compare against. That is Q6's proposal, and this is a concrete confirmation that only the catalog half is missing. Evidence: audits/update-night-2026-09-21/23-B8-floating-pin.txt.
| R-626 | [P2-MEDIUM] An app the customer REMOVED came back: the removal returned 200 and deleted the record and the volume, a container was created two seconds later, and Docker's restart policy has kept it running ever since — while the box reports the app as not installed. FOUND 2026-09-21 during the update night's TEARDOWN, which is the only reason it was found at all. navidrome was removed through the product: the remove_hdd_data:true call was correctly refused 409 (R-442's fail-closed guard — the drive path could not be resolved on this guest), and the remove_hdd_data:false call returned 200 with volumes_removed: ['navidrome_navidrome_data']. Two seconds later a container carrying com.docker.compose.project=navidrome was CREATED (.Created = 19:12:20Z), and Docker's restart: unless-stopped started it again at the next guest boot (.StartedAt = 20:19:33Z, the B5 power cut). Seven hours later: app.yaml absent, deployed=false, one volume back, and the controller happily probing it — Health probe navidrome: API GET :4533/ping → 200. The customer-visible shape is the bad one: "I deleted that app and it came back" — and it came back blank, because the volume really was deleted, so it looks installed and is empty. It is also invisible to every sweep that keys on deployed, which is exactly why the teardown found it and nothing else did. WHAT IS NOT ESTABLISHED, and is stated rather than guessed: what created the container. The controller was restarted several times later in the night and its log no longer reaches that moment — the second time in one night that a restart destroyed the evidence of the thing that mattered (see R-621). Needs: reproduce with a loop that removes an app and watches docker events for 60 s, so the creating path is a NAME and not an inference; then a test that removes an app, reboots, and asserts no container with that compose project exists. And one instrument lesson worth keeping: this session's own post-remove check queried the compose-project label and reported clean at 21:12:18 — two seconds before the container appeared. A check that runs once, immediately, cannot see a thing that is created immediately after it. Evidence: audits/update-night-2026-09-21/26-removed-app-came-back.txt. | READY — rank P2-MEDIUM; owner: CC (controller) |