Files
felhom.eu/documentation/audits/update-night-2026-09-21/MORNING-NOTE.md
T
admin da20722e76
gates / gates (push) Successful in 27s
Update night 2026-09-21: Phase 0 and Phase 1 evidence, the drill method, and two instrument fixes
INTERIM CHECKPOINT — evidence off the machine at the end of the phase that produced it (R-320),
not at the end of the session. Phases 2-5 follow in a later commit.

Phase 0, all three mechanisms proven with their controls:
- the fleet floor to 0.261.0 with its declared MinAgent — both demo boxes in 13 s, the hub
  logging `managed floor SERVED ... from declared (golden 0.258.0)`.
- a PRIVATE DRILL CATALOG (admin/app-catalog-drill), so that broken, dummy, cross-repo and
  engine-major edges can be measured without the live catalog ever carrying one. Positive
  control quoted, and two negative controls: the live catalog's main and both real boxes'
  caches unchanged.
- a throwaway image store on the scratch guest, which is what makes an UNATTENDED HOLD
  measurable at all: an edge that PASSES the within-a-major test and still fails.
  CompareImageRefs was proven to order host:port/ references by RUNNING it (4 positive cases
  + 1 negative control), not by reading it.

Phase 1: real within-a-major upstream edges walked on guest 9202 through the product's own
guarded Update, each app seeded and read back through its OWN front door (R-156), with a
per-edge verdict record in 09's shape. `inconclusive` is never collapsed into `failed`.

TWO INSTRUMENT FIXES, both in this repo's own evidence code:
- 00-api-recipe.md said the app page is /app/<n>; it is /apps/<n>, and every call it described
  404s. Corrected, with the session-expiry note that cost the same time.
- unattended-caller.py's follow() read update_phase/updating off the API ENVELOPE, so both were
  always None and EVERY followed update ran to its 900 s timeout and was then recorded
  `timeout` and never-press-again. Fixed before B1 relied on it. R-623.

No controller, agent or hub code was written. The live catalog carries no broken reference.

Gates: repo_gates.py --fast — all 15 OK, exit 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 21:17:46 +02:00

3.4 KiB

Morning note — to go on top of STATUS.md

(Draft. Numbers marked <…> are filled from the verdict records at the end of the run.)


Updated 2026-09-21 (overnight) — I tested the "update my app" button on as many apps as fit in a night, on good days and bad ones.

updates are proven safe. One is proven dangerous, and the machine handled it exactly right.

Decisions I took on my own: none. Nothing tonight needed a choice you had not already made.

The fleet version is now 0.261.0. You asked for that. Both demo machines took it thirteen seconds after I saved it. The third machine is switched off and will take it when it comes back.

What I exercised. real app updates, each one installed at the version our catalog has today, filled with real data through the app's own front door, backed up, updated to the newer version that really exists upstream, and then the data read back. Plus the two database engines, and the bad days.

The one that matters most. Adventurelog's newer version rewrites the customer's database and then never starts. The machine did everything right: it took a backup one minute before, waited the full five minutes, stopped the app so its data could not be damaged, and told the household in one sentence where the copy is, when it was made, and what is inside it. Then I pressed the restore the sentence names, and the app came back. That app must not be moved to the newer version.

What broke, and whether the household could get out.

Faults found, all written down, none fixed tonight.

  1. Two apps tell the household they are broken when they are fine. Tandoor's health check asks the wrong port; Zipline's asks a web address that app does not have — and the same file already contains the right answer in both cases. A cheap check would catch both.
  2. When an update fails, the machine deletes the broken app's log before anyone can read it. You get "it did not come up" and nothing else. Nobody can say why, or report it upstream.
  3. Pointing a machine at a different app catalog does not work — it keeps reading the old one and says everything is fine. Only matters when you need it, which is the problem.
  4. Smaller ones: a setting the machine calls optional and then refuses without; the catalog password stored in plain text inside the machine's own copy of the catalog.

Rows opened and closed.

What needs you.

  1. Rotate the Gitea admin token. It was printed into my session log by an ordinary git remote -v — the machine stores it in plain text. If you do nothing: the token keeps working and anyone with my session transcript has it.
  2. The promotion list — updates proven safe enough to move on the real catalog. Moving a version is your call, never mine. If you do nothing: nothing breaks; those apps stay where they are and drift further from upstream each month.
  3. The seven open questions about automatic updates now have more facts beside them, and they are still yours. If you do nothing: the automatic-update work cannot start — every part of it hangs off the first question, which is simply may the machine update apps by itself at night?

The live catalog was not touched at all tonight. Not once, not for thirteen minutes. Everything ran against a private drill copy on a scratch machine, exactly as you asked.