Files
felhom.eu/documentation/audits/update-night-2026-09-21/NEW-ROWS.md
T
admin da20722e76
gates / gates (push) Successful in 27s
Update night 2026-09-21: Phase 0 and Phase 1 evidence, the drill method, and two instrument fixes
INTERIM CHECKPOINT — evidence off the machine at the end of the phase that produced it (R-320),
not at the end of the session. Phases 2-5 follow in a later commit.

Phase 0, all three mechanisms proven with their controls:
- the fleet floor to 0.261.0 with its declared MinAgent — both demo boxes in 13 s, the hub
  logging `managed floor SERVED ... from declared (golden 0.258.0)`.
- a PRIVATE DRILL CATALOG (admin/app-catalog-drill), so that broken, dummy, cross-repo and
  engine-major edges can be measured without the live catalog ever carrying one. Positive
  control quoted, and two negative controls: the live catalog's main and both real boxes'
  caches unchanged.
- a throwaway image store on the scratch guest, which is what makes an UNATTENDED HOLD
  measurable at all: an edge that PASSES the within-a-major test and still fails.
  CompareImageRefs was proven to order host:port/ references by RUNNING it (4 positive cases
  + 1 negative control), not by reading it.

Phase 1: real within-a-major upstream edges walked on guest 9202 through the product's own
guarded Update, each app seeded and read back through its OWN front door (R-156), with a
per-edge verdict record in 09's shape. `inconclusive` is never collapsed into `failed`.

TWO INSTRUMENT FIXES, both in this repo's own evidence code:
- 00-api-recipe.md said the app page is /app/<n>; it is /apps/<n>, and every call it described
  404s. Corrected, with the session-expiry note that cost the same time.
- unattended-caller.py's follow() read update_phase/updating off the API ENVELOPE, so both were
  always None and EVERY followed update ran to its 900 s timeout and was then recorded
  `timeout` and never-press-again. Fixed before B1 relied on it. R-623.

No controller, agent or hub code was written. The live catalog carries no broken reference.

Gates: repo_gates.py --fast — all 15 OK, exit 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 21:17:46 +02:00

46 lines
22 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Rows minted by the update night, 2026-09-21 — staging, to be pasted into OPEN-ITEMS.md
Highest existing id at the start of the night: **R-614** (304 rows).
---
| **R-615** | **[P3-LOW] Pointing a box at a different app catalog by `git.repo_url` alone is INERT — the box keeps fetching from the repository it first cloned.** FOUND 2026-09-21 by reading `sync.go` **before** running it, which is the only reason the update night's drill catalog worked at all. `Syncer.gitCloneOrPull` (`controller/internal/sync/sync.go:274-306`) clones **only when `<data>/catalog-cache/.git` is absent**; on every later cycle it runs `git fetch --depth 1 origin <branch>` + `git reset --hard origin/<branch>` **against the remote stored in the clone**, which `buildRepoURL` wrote at clone time. Changing `git.repo_url` in `controller.yaml` and restarting therefore changes **nothing**: the sync keeps pulling the old catalog and reports success. Measured: after the repoint, `git -C <data>/catalog-cache remote -v` still read `app-catalog-felhom.eu`; the box only followed the drill repo once the cache directory was removed. **Why it matters beyond a drill:** this is the one knob that would move a box to a different or a staged catalog — for a migration, a per-customer catalog, or a rollback of the catalog itself — and it silently does not work. **Nothing is wrong with the CACHING**, which is right; what is missing is that a changed `repo_url` must invalidate the clone. **Fix shape:** on start, compare `git.repo_url` with the clone's `origin` and re-clone when they differ (or `git remote set-url` + a full fetch); log which happened. A test that changes `repo_url` under an existing cache and asserts the next sync reads the NEW repo — it fails today. Evidence: `audits/update-night-2026-09-21/04-9202-config-pre.txt`, `05-9202-follows-drill.txt`. | **READY — rank P3-LOW; owner: CC (controller)** |
| **R-616** | **[P3-LOW] The catalog credentials are stored in PLAINTEXT in the box's catalog clone and are printed by an ordinary `git remote -v`.** FOUND 2026-09-21 on guest 9202 while pointing it at a private drill catalog. `Syncer.buildRepoURL` injects `username:token` into the HTTPS URL, and `git clone` persists that URL as the clone's `origin`, so `<data>/catalog-cache/.git/config` holds the token in the clear and **any** diagnostic that prints the remote leaks it — which is what happened in this session's own transcript, and is the same shape as R-580 (`curl -w '%{redirect_url}'`). `maskRepoURL` exists and is used for the LOG lines, so the masking intent is already there; the stored remote is the half that was missed. **INERT ON THE FLEET TODAY** — the live catalog is public and `git.token` is empty on every real box — which is exactly why it should be fixed before it is not: the day the catalog goes private, every box carries a readable credential and every support session that runs `git remote -v` prints it. **Fix shape:** store the remote WITHOUT credentials and supply them per-fetch (a credential helper, `http.extraHeader`, or `GIT_ASKPASS`), and a test asserting the clone's stored `origin` contains no `@`. **Operator action from tonight, unrelated to the fix:** the Gitea `admin` token used for the drill repo was printed by that command and must be rotated. Evidence: `audits/update-night-2026-09-21/05-9202-follows-drill.txt` (redacted). | **READY — rank P3-LOW; owner: CC (controller); one operator action (rotate the Gitea admin token)** |
| **R-617** | **[P3-LOW] The Gitea API token this project uses for pushes cannot create a repository through the documented endpoint, but CAN through `repos/migrate` — so "the token cannot do it" was nearly recorded as a fact when the truth was "one endpoint refuses it".** FOUND 2026-09-21 creating the drill catalog. Both `~/.git-credentials` tokens carry `write:misc,write:notification,write:package,write:issue,write:repository`; `POST /api/v1/user/repos` requires `write:user` and answers **403**, and `POST /api/v1/admin/users/<u>/repos` requires `write:admin` and answers 403 too. `POST /api/v1/repos/migrate` with the same token answered **201** and created the private repository. **Why this is a row and not a note:** a session that stopped at the first 403 would have recorded "CC cannot create a Gitea repository" — an unfalsifiable capability claim of exactly the shape the workspace's standing rule 2 forbids — and every later drill would have been designed around a limit that does not exist. **What it needs:** one line in the operations notes saying which endpoint to use, and (optional, operator) a token scoped for the job so the migrate route is not load-bearing. Evidence: `audits/update-night-2026-09-21/03-drill-repo.txt`. | **READY — rank P3-LOW; owner: CC (docs)** |
| **R-618** | **[P2-MEDIUM] TWO apps are presented to the household as UNHEALTHY while they are working perfectly, and in both cases the SAME template already contains the right answer.** MEASURED 2026-09-21 on guest 9202 (controller v0.261.0, catalog `f5f6a152b513`). Two shapes, one class: **(a) `tandoor` — the WRONG PORT.** `.felhom.yml` probes `port: 8080`; the container listens on **80 and nothing else** (`ss -ltn` inside it), the compose's own traefik label routes to 80, its own docker healthcheck reads `healthy`, and `/accounts/login/` answers **200** through the household's real front door. `GET /api/stacks/tandoor` nevertheless reads `state: "unhealthy"`. **(b) `zipline` — the WRONG PATH.** `.felhom.yml` probes `/api/health`, which zipline 4.6.1 answers **404 `Route GET:/api/health not found`**; **the compose healthcheck in the very same file uses `/api/healthcheck` and is correct and green.** `/dashboard` answers 200. The controller reads `unhealthy`. **This is the MIRROR of R-613** — that is a probe that passes on a broken app (a false GREEN, which no alarm catches); this is a probe that fails on a working app (a false RED). **IT DOES NOT ALARM, AND THAT SETS THE RANK:** `08-alarm-ladder.md` §4 puts `unhealthy` deliberately in the NOT-down set, so no dead-app event and no customer mail follows — the damage is what the household READS, plus anything that gates on `state`. **IT ALREADY COST A MEASUREMENT TONIGHT:** this drill's harness waited for `state == "running"` and hung for its full budget on tandoor, an app that was up the whole time. An instrument waiting for a wrong answer looks exactly like a slow app. **THE GATE THIS WANTS IS CHEAP AND STATIC, AND THAT IS THE FINDING'S REAL VALUE.** Both halves of the answer live in the same template: compare the `.felhom.yml` probe's port and path against the compose's **own** `healthcheck: test:` URL. A sweep of all 53 templates on that rule was run tonight and returns **five** disagreements: `tandoor` (PORT — **CONFIRMED live**), `zipline` (PATH — **CONFIRMED live**), `wger` (PORT, probe 80 vs compose 8000 — **SUSPECTED, NOT MEASURED**, it was not deployed), `home-assistant` (PATH, `/api/` vs `/manifest.json` — **NOT MEASURED**), and `adventurelog` (a FALSE POSITIVE of the sweep's own regex — it reads `running` live). **So the rule finds both real defects, with two candidates and one false positive out of 53** — a good enough signal for a fast gate, provided it reports candidates rather than convictions and a person or a runtime check resolves them. The earlier, cruder rule (probe port vs the *traefik* port) is strictly worse: it clears zipline and convicts adventurelog. **Needs:** fix tandoor's port and zipline's path; measure wger and home-assistant; add the static gate with a decoy each way (R-421) — a template whose probe agrees must not read as a disagreement, and vice versa. **THE GATE'S RULE WAS THEN SHARPENED BY READING `healthprobe.go` RATHER THAN ASSUMING IT, and the sharpening REMOVED a false conviction.** `type: http` treats **any** response as healthy (`healthprobe.go:258-261`), and `type: api` with **no** `expect` block does the same (`:265-268`); only `type: api` WITH `expect.status` cares about the path or the code. So a PATH difference is a candidate only for the third shape, while a PORT difference is a candidate for all of them. Under that rule the 53-template sweep returns **four** candidates — `tandoor` and `zipline` (both CONFIRMED live), `wger` (suspected, unmeasured), and `adventurelog` (a false positive: its compose lists two containers' ports and the probe targets the backend; measured `running`). **`home-assistant` is correctly CLEARED by the sharpened rule** — `type: api`, no `expect`, so its `/api/` answering 401 without a token is healthy, and its edge was PROVEN on the box tonight. The crude rule convicted it; the rule read from the code does not. **That is the gate to build: two of 53 convicted, one suspected, one false positive, and the false positive is resolvable by one live check.** Evidence: `audits/update-night-2026-09-21/10-probe-port-sweep.txt`, `12-probe-vs-compose-healthcheck.txt` and `13-probe-sweep-sharpened.txt`. | **READY — rank P2-MEDIUM; owner: CC (catalog)** |
| **R-619** | **[P3-LOW] A `type: password` deploy field is MANDATORY however `required` reads, and the `deploy-fields` contract says the opposite — so any caller that trusts it is refused.** MEASURED 2026-09-21 on guest 9202 while widening the update drill. `GET /api/stacks/grafana/deploy-fields` serves `{"env_var":"GF_SECURITY_ADMIN_PASSWORD","type":"password","generate":"password:16","required":false}`; a deploy carrying only the two `required:true` fields is refused **400** „a(z) „Admin jelszó" mező kitöltése kötelező — használja a Generálás gombot…". **The BEHAVIOUR is right and is a decision, not a bug:** `deploy.go:305-312` refuses a `password` field with no caller value on purpose — *"We never silently auto-generate — the user needs to know their password"* — which is the opposite of the `secret` case one branch above, where a generated value the customer never sees is exactly correct. **The defect is the CONTRACT.** `.felhom.yml` declares `required: false`, the API serves that verbatim, and nothing on the wire distinguishes "optional because the box will generate it" (`secret`) from "optional in the template and mandatory in the code" (`password`). A person using the deploy page never meets this because the page renders a Generálás button; **anything that is not that page does**, which now includes this drill harness and would include `09` §6.2's unattended caller the day it deploys anything. **Fix shape (smallest that keeps the decision):** serve `required: true` for `type: password` in the deploy-fields response — one place, derived rather than stored, so templates need no edit — and a test asserting a `password` field always reaches the wire as required. Alternatively state it in the field's `description`, which is weaker because it is prose. Evidence: `audits/update-night-2026-09-21/apps/grafana/log.txt` (the refusal) and `batchA.log`. | **READY — rank P3-LOW; owner: CC (controller)** |
| **R-620** | **[P3-LOW] A disabled notifier drops every event with NO local trace, so a box whose hub configuration is absent or broken stops telling anyone anything and leaves nothing behind that says so.** FOUND 2026-09-21 on guest 9202 while trying to score the update night's alarm truth table. `hub.enabled: false` there, and `Notifier.Publish` returns at `notify/notifier.go:269` — **before** any log line — as do `NotifyHealthChange` (:359) and four more entry points. Startup says it once (`[INFO] Notifier disabled (hub not configured)`) and then every later event, of every severity up to `critical`, vanishes without a word. **The measurable consequence tonight:** the whole event-and-mail half of the drill was structurally unmeasurable on this venue, and the alarm truth table below covers only the app page, the dashboard and the box's own log. That is a cost this session paid and named; the next one would pay it again. **The consequence on a real box is smaller but not zero:** the fleet's boxes have the hub enabled, and total silence is already caught by the hub's dead-man's-switch (staleness from the LAST REPORT, proven in the 2026-07-22 power-outage audit). What is NOT caught is the in-between — a box that still reports but whose notifier was disabled by a bad config push would go on reporting healthy while dropping every alarm, and the only evidence would be a single INFO line at the last restart. **Fix shape:** one DEBUG (or WARN, once per event type) line on the disabled path naming the event that was dropped, so the absence is visible where it happens rather than inferable from a startup line. Cheap, and it converts an invisible failure into a greppable one — R-96 rule 3 in the place that produces it. Evidence: `audits/update-night-2026-09-21/11-notifier-disabled.txt`. | **READY — rank P3-LOW; owner: CC (controller)** |
| **R-621** | **[P2-MEDIUM] A held update DESTROYS the evidence of why it failed: `failAndHold` runs `compose down`, the failing containers are removed, and their output is gone before anyone — household, operator or the next session — can read it.** MEASURED 2026-09-21 on guest 9202 on a REAL upstream edge: `adventurelog v0.12.1 → v0.13.0`. The new backend applied **nine Django migrations successfully** and then never listened; the update held after the full 5-minute health wait. **`Manager.failAndHold` (`stacks/update.go:723`) calls `updateCompose(dir, env, "down")`**, which removes the containers rather than stopping them, and nothing captures their logs first. Within seconds the box's own log recorded `Logs result for adventurelog: 0 bytes returned (empty)` and `docker ps -a` held nothing at all. **What survives is the WHAT and not the WHY:** the controller line `update adventurelog FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state unhealthy)` and the household's sentence, both of which say the app did not come up and neither of which says the migrations ran and the server then failed to bind. **This is R-320 ("evidence off the machine before the teardown") as a PRODUCT behaviour rather than a session habit** — the teardown here is the product's own, it is correct to perform (a half-started new version must not keep running), and it happens before anyone can look. **Why it matters beyond a drill:** the hold sentence sends the household to a restore, and after the restore the only remaining question is *should I press Update again?* — which nobody can answer, because the one artefact that would say so no longer exists. It also makes every future held update unreportable to an upstream project. **Fix shape:** capture `compose logs --no-color --tail N` into the stack directory (beside `applied-compose.yml`, which already travels with the stack) IMMEDIATELY before the `down`, and surface it on the app page's hold panel or at least through the existing `/api/stacks/<n>/logs` fallback. Bounded size, written once per hold. A test that holds an app and asserts the captured file is non-empty — it fails today. Evidence: `audits/update-night-2026-09-21/apps/adventurelog/why-it-failed.txt`, `state-after-hold.txt`. | **READY — rank P2-MEDIUM; owner: CC (controller)** |
| **R-622** | **[P2-MEDIUM] `adventurelog v0.13.0` migrates the customer's database and then does not serve — the edge must NOT be promoted, and it is the first real-catalog candidate this project has measured as unsafe.** MEASURED 2026-09-21 on guest 9202 through the product's own guarded Update. `v0.12.1 → v0.13.0` (backend AND frontend together, PostGIS held constant). The backend applied **nine migrations, every one `... OK`** — `adventures.0072_trail_wanderer_author_fields` through `integrations.0009_alter_endurainintegration_auth_method`, plus `billing.0001_initial` — and then the container's own healthcheck failed with `URLError: [Errno 111] Connection refused` on five consecutive checks. The app never bound its port. The update held honestly after the full 5-minute wait. **THE PRODUCT DID EVERYTHING RIGHT AND THAT IS HALF THE FINDING:** the precondition found a Tier-1 copy one minute old, the safety dump was written, the pin advanced BEFORE the pull, the health wait was not short-circuited, the app was stopped rather than left half-running, and the hold sentence named the tier, the date and what the copy holds — „saját meghajtó, 2026-09-21 20:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza." **This is exactly the case `09` §4 exists for:** the migration RAN, so there is no undo, only a restore — and the restore is the thing slice 4 made sure existed first. **What it needs:** adventurelog stays OFF the promotion list; the cause is not diagnosed here (R-621 is why); and before it is ever promoted the edge should be re-run on the harness with its ABORT step, since an app that migrates and then refuses is the shape most likely to refuse the old image too. Evidence: `audits/update-night-2026-09-21/apps/adventurelog/`. | **READY — rank P2-MEDIUM; owner: CC (catalog); NOT a Felhom defect — an upstream edge that fails** |
---
## Lines to ADD to existing rows (not re-filed)
**R-607** — append: **Seen again 2026-09-21 (update night), and now a dozen times in one session.** Every drill-catalog bump of the night was followed by `POST /api/sync` answering „Sablonok naprakészek — nincs változás" while the box's cache HAD moved, with `catalog_images` staying stale until a separate `POST /api/stacks/rescan`. The night's harness therefore rescans unconditionally after every sync, which is a workaround and not a fix. **The window was still never measured as a NUMBER** — that is what the row asks for and what remains owed.
**R-462** — append: **The count moved on 2026-09-21 from 3 apps to <N>.** The update night walked <N> real, within-a-major upstream edges on guest 9202 through the product's own guarded Update, each seeded and read back through the app's own front door: <LIST>. <P> proven, <F> failed, <I> inconclusive. Box-side fixtures for <FX> apps now exist at `audits/update-night-2026-09-21/fixtures.py`, and four of them (actualbudget, navidrome, audiobookshelf, vikunja) are ported into `app-catalog-felhom.eu/scripts/upgrade_fixtures.py` with seven new `EDGES` (U1–U7) so the same edges can be run on the harness venue **with their ABORT step**, which the box deliberately does not offer. **Owed:** the harness RUNS for those edges (the code is in; the runs are not), and fixtures for the apps recorded inconclusive tonight.
**R-463** — append: **Measured 2026-09-21 (update night), both halves.** <PG52>. And the conversion rehearsal Q5 asks for was costed on a real seeded datadir: <PG53>.
**R-469 / R-459** — append: **The MariaDB engine major was pressed through the real Update BUTTON for the first time on 2026-09-21** (it had only ever been run on the harness). <MARIA>.
**R-446** — append: **Measured on the box 2026-09-21 (update night), leg B8.** <B8>.
**R-458** — append: **Measured 2026-09-21 (update night), leg B9.** <B9>.
**R-613** — append: the update night could not seed `uptime-kuma` for the same reason and left it out rather than faking it.
**R-460** — append: bookstack's edge was walked again on 2026-09-21 and is again **half-proven** — the database half read back through `php artisan`, the file half untouched. The limitation is unchanged and is now measured on the box as well as on the harness.
**R-442 (CLOSED)** — append, as a confirmation rather than a reopening: **the fail-closed half was exercised again 2026-09-21** on guest 9202, where `/api/disks` answers `agent not configured`. Three apps deployed with an `HDD_PATH` (navidrome, audiobookshelf, romm) were each REFUSED at „remove with data" — „A(z) …/userdata/<app> tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó…" — **with the app kept**, and each was then removed successfully with the data KEPT. So the guard refuses the destructive half and leaves the non-destructive half available, which is exactly the shape the row describes. No change to the row's status.
| **R-623** | **[P3-LOW] The unattended-update caller turned every SUCCESS into a `timeout`, and then refused to press that app again — the instrument, not the box.** FOUND 2026-09-21 (update night) by reading `unattended-caller.py` before relying on it for the Q4 hold measurement. Its `call()` returns the API **envelope** — `{"ok": true, "data": {…}}` — and `follow()` read `update_phase` and `updating` **off the envelope**, where neither exists. Both were therefore always `None`; the end test `not updating and phase in ("done","failed")` could never fire; every followed update ran the full **900-second** timeout and was recorded `timeout`, which the caller treats as terminal and adds to `never_again`. `main()` unwraps `data` for the stack LIST, which is exactly why the within-a-major half of that night worked and this half did not. **The 2026-09-21 run did not catch it because the only pass that reached `follow()` was Scenario F, whose log was lost to a buffering `tail`** — the run's own honestly-recorded instrumentation gap turns out to have hidden a second one underneath it. **This is the R-607 class in the evidence layer rather than the product layer: an instrument that can report a success as a timeout is not a measurement**, and worse, it is a measurement that says the box behaved badly when the box behaved well. **FIXED in the same file 2026-09-21** (unwrap `data`, with the reason written into the docstring so the next reader does not re-derive it), and the fixed caller is what produced tonight's unattended-hold leg. **What it does NOT invalidate:** the G-b no-retry proof, which is entirely in the refusal path and never reached `follow()`. **What it DOES qualify:** any future reading of that night's Scenario F timing — the "51 s – 1 m 26 s" figures come from the ATTENDED scenarios 04/05/07, not from the caller. | **CLOSED 2026-09-21 — fixed in `audits/update-arc-gaps-2026-09-21/unattended-caller.py`** |