Files
felhom.eu/documentation/audits/update-night-2026-09-21/NEW-ROWS.md
T
admin da20722e76
gates / gates (push) Successful in 27s
Update night 2026-09-21: Phase 0 and Phase 1 evidence, the drill method, and two instrument fixes
INTERIM CHECKPOINT — evidence off the machine at the end of the phase that produced it (R-320),
not at the end of the session. Phases 2-5 follow in a later commit.

Phase 0, all three mechanisms proven with their controls:
- the fleet floor to 0.261.0 with its declared MinAgent — both demo boxes in 13 s, the hub
  logging `managed floor SERVED ... from declared (golden 0.258.0)`.
- a PRIVATE DRILL CATALOG (admin/app-catalog-drill), so that broken, dummy, cross-repo and
  engine-major edges can be measured without the live catalog ever carrying one. Positive
  control quoted, and two negative controls: the live catalog's main and both real boxes'
  caches unchanged.
- a throwaway image store on the scratch guest, which is what makes an UNATTENDED HOLD
  measurable at all: an edge that PASSES the within-a-major test and still fails.
  CompareImageRefs was proven to order host:port/ references by RUNNING it (4 positive cases
  + 1 negative control), not by reading it.

Phase 1: real within-a-major upstream edges walked on guest 9202 through the product's own
guarded Update, each app seeded and read back through its OWN front door (R-156), with a
per-edge verdict record in 09's shape. `inconclusive` is never collapsed into `failed`.

TWO INSTRUMENT FIXES, both in this repo's own evidence code:
- 00-api-recipe.md said the app page is /app/<n>; it is /apps/<n>, and every call it described
  404s. Corrected, with the session-expiry note that cost the same time.
- unattended-caller.py's follow() read update_phase/updating off the API ENVELOPE, so both were
  always None and EVERY followed update ran to its 900 s timeout and was then recorded
  `timeout` and never-press-again. Fixed before B1 relied on it. R-623.

No controller, agent or hub code was written. The live catalog carries no broken reference.

Gates: repo_gates.py --fast — all 15 OK, exit 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 21:17:46 +02:00

22 KiB
Raw Blame History

Rows minted by the update night, 2026-09-21 — staging, to be pasted into OPEN-ITEMS.md

Highest existing id at the start of the night: R-614 (304 rows).


| R-615 | [P3-LOW] Pointing a box at a different app catalog by git.repo_url alone is INERT — the box keeps fetching from the repository it first cloned. FOUND 2026-09-21 by reading sync.go before running it, which is the only reason the update night's drill catalog worked at all. Syncer.gitCloneOrPull (controller/internal/sync/sync.go:274-306) clones only when <data>/catalog-cache/.git is absent; on every later cycle it runs git fetch --depth 1 origin <branch> + git reset --hard origin/<branch> against the remote stored in the clone, which buildRepoURL wrote at clone time. Changing git.repo_url in controller.yaml and restarting therefore changes nothing: the sync keeps pulling the old catalog and reports success. Measured: after the repoint, git -C <data>/catalog-cache remote -v still read app-catalog-felhom.eu; the box only followed the drill repo once the cache directory was removed. Why it matters beyond a drill: this is the one knob that would move a box to a different or a staged catalog — for a migration, a per-customer catalog, or a rollback of the catalog itself — and it silently does not work. Nothing is wrong with the CACHING, which is right; what is missing is that a changed repo_url must invalidate the clone. Fix shape: on start, compare git.repo_url with the clone's origin and re-clone when they differ (or git remote set-url + a full fetch); log which happened. A test that changes repo_url under an existing cache and asserts the next sync reads the NEW repo — it fails today. Evidence: audits/update-night-2026-09-21/04-9202-config-pre.txt, 05-9202-follows-drill.txt. | READY — rank P3-LOW; owner: CC (controller) |

| R-616 | [P3-LOW] The catalog credentials are stored in PLAINTEXT in the box's catalog clone and are printed by an ordinary git remote -v. FOUND 2026-09-21 on guest 9202 while pointing it at a private drill catalog. Syncer.buildRepoURL injects username:token into the HTTPS URL, and git clone persists that URL as the clone's origin, so <data>/catalog-cache/.git/config holds the token in the clear and any diagnostic that prints the remote leaks it — which is what happened in this session's own transcript, and is the same shape as R-580 (curl -w '%{redirect_url}'). maskRepoURL exists and is used for the LOG lines, so the masking intent is already there; the stored remote is the half that was missed. INERT ON THE FLEET TODAY — the live catalog is public and git.token is empty on every real box — which is exactly why it should be fixed before it is not: the day the catalog goes private, every box carries a readable credential and every support session that runs git remote -v prints it. Fix shape: store the remote WITHOUT credentials and supply them per-fetch (a credential helper, http.extraHeader, or GIT_ASKPASS), and a test asserting the clone's stored origin contains no @. Operator action from tonight, unrelated to the fix: the Gitea admin token used for the drill repo was printed by that command and must be rotated. Evidence: audits/update-night-2026-09-21/05-9202-follows-drill.txt (redacted). | READY — rank P3-LOW; owner: CC (controller); one operator action (rotate the Gitea admin token) |

| R-617 | [P3-LOW] The Gitea API token this project uses for pushes cannot create a repository through the documented endpoint, but CAN through repos/migrate — so "the token cannot do it" was nearly recorded as a fact when the truth was "one endpoint refuses it". FOUND 2026-09-21 creating the drill catalog. Both ~/.git-credentials tokens carry write:misc,write:notification,write:package,write:issue,write:repository; POST /api/v1/user/repos requires write:user and answers 403, and POST /api/v1/admin/users/<u>/repos requires write:admin and answers 403 too. POST /api/v1/repos/migrate with the same token answered 201 and created the private repository. Why this is a row and not a note: a session that stopped at the first 403 would have recorded "CC cannot create a Gitea repository" — an unfalsifiable capability claim of exactly the shape the workspace's standing rule 2 forbids — and every later drill would have been designed around a limit that does not exist. What it needs: one line in the operations notes saying which endpoint to use, and (optional, operator) a token scoped for the job so the migrate route is not load-bearing. Evidence: audits/update-night-2026-09-21/03-drill-repo.txt. | READY — rank P3-LOW; owner: CC (docs) |

| R-618 | [P2-MEDIUM] TWO apps are presented to the household as UNHEALTHY while they are working perfectly, and in both cases the SAME template already contains the right answer. MEASURED 2026-09-21 on guest 9202 (controller v0.261.0, catalog f5f6a152b513). Two shapes, one class: (a) tandoor — the WRONG PORT. .felhom.yml probes port: 8080; the container listens on 80 and nothing else (ss -ltn inside it), the compose's own traefik label routes to 80, its own docker healthcheck reads healthy, and /accounts/login/ answers 200 through the household's real front door. GET /api/stacks/tandoor nevertheless reads state: "unhealthy". (b) zipline — the WRONG PATH. .felhom.yml probes /api/health, which zipline 4.6.1 answers 404 Route GET:/api/health not found; the compose healthcheck in the very same file uses /api/healthcheck and is correct and green. /dashboard answers 200. The controller reads unhealthy. This is the MIRROR of R-613 — that is a probe that passes on a broken app (a false GREEN, which no alarm catches); this is a probe that fails on a working app (a false RED). IT DOES NOT ALARM, AND THAT SETS THE RANK: 08-alarm-ladder.md §4 puts unhealthy deliberately in the NOT-down set, so no dead-app event and no customer mail follows — the damage is what the household READS, plus anything that gates on state. IT ALREADY COST A MEASUREMENT TONIGHT: this drill's harness waited for state == "running" and hung for its full budget on tandoor, an app that was up the whole time. An instrument waiting for a wrong answer looks exactly like a slow app. THE GATE THIS WANTS IS CHEAP AND STATIC, AND THAT IS THE FINDING'S REAL VALUE. Both halves of the answer live in the same template: compare the .felhom.yml probe's port and path against the compose's own healthcheck: test: URL. A sweep of all 53 templates on that rule was run tonight and returns five disagreements: tandoor (PORT — CONFIRMED live), zipline (PATH — CONFIRMED live), wger (PORT, probe 80 vs compose 8000 — SUSPECTED, NOT MEASURED, it was not deployed), home-assistant (PATH, /api/ vs /manifest.json — NOT MEASURED), and adventurelog (a FALSE POSITIVE of the sweep's own regex — it reads running live). So the rule finds both real defects, with two candidates and one false positive out of 53 — a good enough signal for a fast gate, provided it reports candidates rather than convictions and a person or a runtime check resolves them. The earlier, cruder rule (probe port vs the traefik port) is strictly worse: it clears zipline and convicts adventurelog. Needs: fix tandoor's port and zipline's path; measure wger and home-assistant; add the static gate with a decoy each way (R-421) — a template whose probe agrees must not read as a disagreement, and vice versa. THE GATE'S RULE WAS THEN SHARPENED BY READING healthprobe.go RATHER THAN ASSUMING IT, and the sharpening REMOVED a false conviction. type: http treats any response as healthy (healthprobe.go:258-261), and type: api with no expect block does the same (:265-268); only type: api WITH expect.status cares about the path or the code. So a PATH difference is a candidate only for the third shape, while a PORT difference is a candidate for all of them. Under that rule the 53-template sweep returns four candidates — tandoor and zipline (both CONFIRMED live), wger (suspected, unmeasured), and adventurelog (a false positive: its compose lists two containers' ports and the probe targets the backend; measured running). home-assistant is correctly CLEARED by the sharpened rule — type: api, no expect, so its /api/ answering 401 without a token is healthy, and its edge was PROVEN on the box tonight. The crude rule convicted it; the rule read from the code does not. That is the gate to build: two of 53 convicted, one suspected, one false positive, and the false positive is resolvable by one live check. Evidence: audits/update-night-2026-09-21/10-probe-port-sweep.txt, 12-probe-vs-compose-healthcheck.txt and 13-probe-sweep-sharpened.txt. | READY — rank P2-MEDIUM; owner: CC (catalog) |

| R-619 | [P3-LOW] A type: password deploy field is MANDATORY however required reads, and the deploy-fields contract says the opposite — so any caller that trusts it is refused. MEASURED 2026-09-21 on guest 9202 while widening the update drill. GET /api/stacks/grafana/deploy-fields serves {"env_var":"GF_SECURITY_ADMIN_PASSWORD","type":"password","generate":"password:16","required":false}; a deploy carrying only the two required:true fields is refused 400 „a(z) „Admin jelszó" mező kitöltése kötelező — használja a Generálás gombot…". The BEHAVIOUR is right and is a decision, not a bug: deploy.go:305-312 refuses a password field with no caller value on purpose — "We never silently auto-generate — the user needs to know their password" — which is the opposite of the secret case one branch above, where a generated value the customer never sees is exactly correct. The defect is the CONTRACT. .felhom.yml declares required: false, the API serves that verbatim, and nothing on the wire distinguishes "optional because the box will generate it" (secret) from "optional in the template and mandatory in the code" (password). A person using the deploy page never meets this because the page renders a Generálás button; anything that is not that page does, which now includes this drill harness and would include 09 §6.2's unattended caller the day it deploys anything. Fix shape (smallest that keeps the decision): serve required: true for type: password in the deploy-fields response — one place, derived rather than stored, so templates need no edit — and a test asserting a password field always reaches the wire as required. Alternatively state it in the field's description, which is weaker because it is prose. Evidence: audits/update-night-2026-09-21/apps/grafana/log.txt (the refusal) and batchA.log. | READY — rank P3-LOW; owner: CC (controller) |

| R-620 | [P3-LOW] A disabled notifier drops every event with NO local trace, so a box whose hub configuration is absent or broken stops telling anyone anything and leaves nothing behind that says so. FOUND 2026-09-21 on guest 9202 while trying to score the update night's alarm truth table. hub.enabled: false there, and Notifier.Publish returns at notify/notifier.go:269 — before any log line — as do NotifyHealthChange (:359) and four more entry points. Startup says it once ([INFO] Notifier disabled (hub not configured)) and then every later event, of every severity up to critical, vanishes without a word. The measurable consequence tonight: the whole event-and-mail half of the drill was structurally unmeasurable on this venue, and the alarm truth table below covers only the app page, the dashboard and the box's own log. That is a cost this session paid and named; the next one would pay it again. The consequence on a real box is smaller but not zero: the fleet's boxes have the hub enabled, and total silence is already caught by the hub's dead-man's-switch (staleness from the LAST REPORT, proven in the 2026-07-22 power-outage audit). What is NOT caught is the in-between — a box that still reports but whose notifier was disabled by a bad config push would go on reporting healthy while dropping every alarm, and the only evidence would be a single INFO line at the last restart. Fix shape: one DEBUG (or WARN, once per event type) line on the disabled path naming the event that was dropped, so the absence is visible where it happens rather than inferable from a startup line. Cheap, and it converts an invisible failure into a greppable one — R-96 rule 3 in the place that produces it. Evidence: audits/update-night-2026-09-21/11-notifier-disabled.txt. | READY — rank P3-LOW; owner: CC (controller) |

| R-621 | [P2-MEDIUM] A held update DESTROYS the evidence of why it failed: failAndHold runs compose down, the failing containers are removed, and their output is gone before anyone — household, operator or the next session — can read it. MEASURED 2026-09-21 on guest 9202 on a REAL upstream edge: adventurelog v0.12.1 → v0.13.0. The new backend applied nine Django migrations successfully and then never listened; the update held after the full 5-minute health wait. Manager.failAndHold (stacks/update.go:723) calls updateCompose(dir, env, "down"), which removes the containers rather than stopping them, and nothing captures their logs first. Within seconds the box's own log recorded Logs result for adventurelog: 0 bytes returned (empty) and docker ps -a held nothing at all. What survives is the WHAT and not the WHY: the controller line update adventurelog FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state unhealthy) and the household's sentence, both of which say the app did not come up and neither of which says the migrations ran and the server then failed to bind. This is R-320 ("evidence off the machine before the teardown") as a PRODUCT behaviour rather than a session habit — the teardown here is the product's own, it is correct to perform (a half-started new version must not keep running), and it happens before anyone can look. Why it matters beyond a drill: the hold sentence sends the household to a restore, and after the restore the only remaining question is should I press Update again? — which nobody can answer, because the one artefact that would say so no longer exists. It also makes every future held update unreportable to an upstream project. Fix shape: capture compose logs --no-color --tail N into the stack directory (beside applied-compose.yml, which already travels with the stack) IMMEDIATELY before the down, and surface it on the app page's hold panel or at least through the existing /api/stacks/<n>/logs fallback. Bounded size, written once per hold. A test that holds an app and asserts the captured file is non-empty — it fails today. Evidence: audits/update-night-2026-09-21/apps/adventurelog/why-it-failed.txt, state-after-hold.txt. | READY — rank P2-MEDIUM; owner: CC (controller) |

| R-622 | [P2-MEDIUM] adventurelog v0.13.0 migrates the customer's database and then does not serve — the edge must NOT be promoted, and it is the first real-catalog candidate this project has measured as unsafe. MEASURED 2026-09-21 on guest 9202 through the product's own guarded Update. v0.12.1 → v0.13.0 (backend AND frontend together, PostGIS held constant). The backend applied nine migrations, every one ... OK — adventures.0072_trail_wanderer_author_fields through integrations.0009_alter_endurainintegration_auth_method, plus billing.0001_initial — and then the container's own healthcheck failed with URLError: [Errno 111] Connection refused on five consecutive checks. The app never bound its port. The update held honestly after the full 5-minute wait. THE PRODUCT DID EVERYTHING RIGHT AND THAT IS HALF THE FINDING: the precondition found a Tier-1 copy one minute old, the safety dump was written, the pin advanced BEFORE the pull, the health wait was not short-circuited, the app was stopped rather than left half-running, and the hold sentence named the tier, the date and what the copy holds — „saját meghajtó, 2026-09-21 20:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza." This is exactly the case 09 §4 exists for: the migration RAN, so there is no undo, only a restore — and the restore is the thing slice 4 made sure existed first. What it needs: adventurelog stays OFF the promotion list; the cause is not diagnosed here (R-621 is why); and before it is ever promoted the edge should be re-run on the harness with its ABORT step, since an app that migrates and then refuses is the shape most likely to refuse the old image too. Evidence: audits/update-night-2026-09-21/apps/adventurelog/. | READY — rank P2-MEDIUM; owner: CC (catalog); NOT a Felhom defect — an upstream edge that fails |


Lines to ADD to existing rows (not re-filed)

R-607 — append: Seen again 2026-09-21 (update night), and now a dozen times in one session. Every drill-catalog bump of the night was followed by POST /api/sync answering „Sablonok naprakészek — nincs változás" while the box's cache HAD moved, with catalog_images staying stale until a separate POST /api/stacks/rescan. The night's harness therefore rescans unconditionally after every sync, which is a workaround and not a fix. The window was still never measured as a NUMBER — that is what the row asks for and what remains owed.

R-462 — append: The count moved on 2026-09-21 from 3 apps to . The update night walked real, within-a-major upstream edges on guest 9202 through the product's own guarded Update, each seeded and read back through the app's own front door: .

proven, failed, inconclusive. Box-side fixtures for apps now exist at audits/update-night-2026-09-21/fixtures.py, and four of them (actualbudget, navidrome, audiobookshelf, vikunja) are ported into app-catalog-felhom.eu/scripts/upgrade_fixtures.py with seven new EDGES (U1–U7) so the same edges can be run on the harness venue with their ABORT step, which the box deliberately does not offer. Owed: the harness RUNS for those edges (the code is in; the runs are not), and fixtures for the apps recorded inconclusive tonight.

R-463 — append: Measured 2026-09-21 (update night), both halves. . And the conversion rehearsal Q5 asks for was costed on a real seeded datadir: .

R-469 / R-459 — append: The MariaDB engine major was pressed through the real Update BUTTON for the first time on 2026-09-21 (it had only ever been run on the harness). .

R-446 — append: Measured on the box 2026-09-21 (update night), leg B8. .

R-458 — append: Measured 2026-09-21 (update night), leg B9. .

R-613 — append: the update night could not seed uptime-kuma for the same reason and left it out rather than faking it.

R-460 — append: bookstack's edge was walked again on 2026-09-21 and is again half-proven — the database half read back through php artisan, the file half untouched. The limitation is unchanged and is now measured on the box as well as on the harness.

R-442 (CLOSED) — append, as a confirmation rather than a reopening: the fail-closed half was exercised again 2026-09-21 on guest 9202, where /api/disks answers agent not configured. Three apps deployed with an HDD_PATH (navidrome, audiobookshelf, romm) were each REFUSED at „remove with data" — „A(z) …/userdata/ tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó…" — with the app kept, and each was then removed successfully with the data KEPT. So the guard refuses the destructive half and leaves the non-destructive half available, which is exactly the shape the row describes. No change to the row's status.

| R-623 | [P3-LOW] The unattended-update caller turned every SUCCESS into a timeout, and then refused to press that app again — the instrument, not the box. FOUND 2026-09-21 (update night) by reading unattended-caller.py before relying on it for the Q4 hold measurement. Its call() returns the API envelope — {"ok": true, "data": {…}} — and follow() read update_phase and updating off the envelope, where neither exists. Both were therefore always None; the end test not updating and phase in ("done","failed") could never fire; every followed update ran the full 900-second timeout and was recorded timeout, which the caller treats as terminal and adds to never_again. main() unwraps data for the stack LIST, which is exactly why the within-a-major half of that night worked and this half did not. The 2026-09-21 run did not catch it because the only pass that reached follow() was Scenario F, whose log was lost to a buffering tail — the run's own honestly-recorded instrumentation gap turns out to have hidden a second one underneath it. This is the R-607 class in the evidence layer rather than the product layer: an instrument that can report a success as a timeout is not a measurement, and worse, it is a measurement that says the box behaved badly when the box behaved well. FIXED in the same file 2026-09-21 (unwrap data, with the reason written into the docstring so the next reader does not re-derive it), and the fixed caller is what produced tonight's unattended-hold leg. What it does NOT invalidate: the G-b no-retry proof, which is entirely in the refusal path and never reached follow(). What it DOES qualify: any future reading of that night's Scenario F timing — the "51 s – 1 m 26 s" figures come from the ATTENDED scenarios 04/05/07, not from the caller. | CLOSED 2026-09-21 — fixed in audits/update-arc-gaps-2026-09-21/unattended-caller.py |