Files
felhom.eu/documentation/audits/update-night-2026-09-21/PROGRESS.md
T
admin da20722e76
gates / gates (push) Successful in 27s
Update night 2026-09-21: Phase 0 and Phase 1 evidence, the drill method, and two instrument fixes
INTERIM CHECKPOINT — evidence off the machine at the end of the phase that produced it (R-320),
not at the end of the session. Phases 2-5 follow in a later commit.

Phase 0, all three mechanisms proven with their controls:
- the fleet floor to 0.261.0 with its declared MinAgent — both demo boxes in 13 s, the hub
  logging `managed floor SERVED ... from declared (golden 0.258.0)`.
- a PRIVATE DRILL CATALOG (admin/app-catalog-drill), so that broken, dummy, cross-repo and
  engine-major edges can be measured without the live catalog ever carrying one. Positive
  control quoted, and two negative controls: the live catalog's main and both real boxes'
  caches unchanged.
- a throwaway image store on the scratch guest, which is what makes an UNATTENDED HOLD
  measurable at all: an edge that PASSES the within-a-major test and still fails.
  CompareImageRefs was proven to order host:port/ references by RUNNING it (4 positive cases
  + 1 negative control), not by reading it.

Phase 1: real within-a-major upstream edges walked on guest 9202 through the product's own
guarded Update, each app seeded and read back through its OWN front door (R-156), with a
per-edge verdict record in 09's shape. `inconclusive` is never collapsed into `failed`.

TWO INSTRUMENT FIXES, both in this repo's own evidence code:
- 00-api-recipe.md said the app page is /app/<n>; it is /apps/<n>, and every call it described
  404s. Corrected, with the session-expiry note that cost the same time.
- unattended-caller.py's follow() read update_phase/updating off the API ENVELOPE, so both were
  always None and EVERY followed update ran to its 900 s timeout and was then recorded
  `timeout` and never-press-again. Fixed before B1 relied on it. R-623.

No controller, agent or hub code was written. The live catalog carries no broken reference.

Gates: repo_gates.py --fast — all 15 OK, exit 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 21:17:46 +02:00

5.6 KiB

UPDATE NIGHT 2026-09-21 — progress log (one line per finished step)

Evidence dir: documentation/audits/update-night-2026-09-21/ A resuming session reads THIS FILE FIRST and never repeats a finished step.

time (CEST) step verdict evidence
20:07 P0.1 floor → 0.261.0 (declared MinAgent 0.131.0) PROVEN — both demo boxes arrived in 13 s; hub managed floor SERVED … from declared (golden 0.258.0); drill-r50 stays held (agent 0.129.0 < 0.131.0, host DOWN) 01-floor-pre.txt, 02-floor-save.txt
20:08 P0.2a drill repo admin/app-catalog-drill created (private) DONE — Gitea token lacks write:user, so POST /api/v1/repos/migrate was used instead of /user/repos; main = f5f6a152b513 = live 03-drill-repo.txt
20:10 P0.2b 9202 repointed at the drill repo PROVEN — git.repo_url ALONE IS INERT: gitCloneOrPull only clones when .git is absent, else fetches from the old origin. Cache dir had to be removed too. Config saved as controller.yaml.pre-update-night 04-9202-config-pre.txt, 05-9202-follows-drill.txt
20:12 P0.2c positive + negative controls PROVEN — drill bump uptime-kuma 2.4.0→2.5.5 shows on 9202 as „Frissítés elérhető — ma" / "Update available — today"; live catalog still f5f6a152b513 with pin 2.4.0; both real boxes' caches still at f5f6a15. R-607 fired again (sync said „nincs változás"). App route is /apps/<n>, not /app/<n> 06-drift-rerun.txt, 07-positive-control.txt
20:07 Drift re-run CONFIRMED — 66 pins, 46 behind, 39 within-major, 7 across-major — the brief's numbers hold exactly 06-drift-rerun.txt
20:13 P0.4 capacity measured OK — 9202: 26 GB RAM (22 free), docker root on mp0 with 56 GB free, drive 872 GB. demo-hp / is 86% but holds neither. Brief's capacity claim HOLDS 08-capacity.txt
20:18 P0.3 throwaway image store PROVEN — registry:2 on 9202 127.0.0.1:5000; drill/glance:1.0.0 = real v0.8.6 retagged, :1.0.1 = starts/stays up/never serves, :1.0.2 absent (404). CompareImageRefs DOES order host:port/ refs — 4 positive + 1 negative control, run not read; tmp test deleted, tree clean 09-image-store.txt
20:20 EDGE privatebin 2.0.5 -> 2.0.6 (file-leg) PROVEN — 15.4 s; seed read back both sides; all four observables agree apps/privatebin/
20:25 EDGE docmost 0.95.0 -> 0.96.0 (db-postgres, engine constant) PROVEN — 103.5 s; seeded account authenticated after; all four observables agree apps/docmost/
20:28 EDGE bookstack 26.05.2 -> 26.05.5 (db-mariadb, engine constant) PROVEN — 45.1 s; artisan readback with its own negative control; DATABASE HALF ONLY (R-460) apps/bookstack/
20:37 FINDING R-618 tandoor probe port DEFECT, measured — probe 8080, app listens on 80 only; docker healthy + front door 200 + controller unhealthy. 53-template sweep run: wger suspected, adventurelog cleared 10-probe-port-sweep.txt
20:38 FINDINGS R-615/616/617/619 filed — repo_url inert; catalog token plaintext in the clone; Gitea migrate endpoint; type: password mandatory though the wire says optional NEW-ROWS.md
20:39 EDGE actualbudget 26.7.0 -> 26.9.0 PROVEN — 19.5 s apps/actualbudget/
20:44 EDGE navidrome 0.63.2 -> 0.64.0 (file-leg, HDD_PATH) PROVEN — 11.3 s apps/navidrome/
20:46 EDGE audiobookshelf 2.35.1 -> 2.36.1 (file-leg) PROVEN — 23.6 s apps/audiobookshelf/
20:54 EDGE adventurelog v0.12.1 -> v0.13.0 (db-postgis) FAILED — the most valuable result so far. Nine migrations applied OK, then the app never bound its port; held after the full 5-min wait; the sentence names tier, date and what the copy holds. Rows R-621 (the hold destroys the failure evidence) and R-622 (do not promote this edge) apps/adventurelog/
20:59 Harness code pushed to the catalog — 4 fixtures + edges U1..U7 DONE, gates green — live catalog main moves f5f6a152b513 -> 4463243f2e09, scripts/ ONLY, zero image: lines (the teardown diff still expects every image line identical). Harness RUNS owed app-catalog-felhom.eu@4463243f2e09
20:59 image reclaim BY NAME (no prune) 34 unused images removed by exact reference; 40 GB -> 55 GB free reclaim.sh
21:08 CI check for the catalog push GREEN — job id=830, name='gates', status='completed', conclusion='success', matched on head_sha=4463243f2e09. Confirms CLAUDE.md's warning: page 1's newest id was 52, the real newest was 830 — job ids are NOT page-ordered scratchpad/ci.sh
21:12 Observation, NOT a new row — an app with an HDD_PATH cannot have its DATA removed on 9202 R-442's guard working as designed: /api/disks answers agent not configured on this guest, so the drive path cannot be RESOLVED and the removal is refused with the app kept rather than half-deleted. The household's other choice — remove the app, KEEP the data — is accepted. The harness now takes that route and tidies its own directories by name at teardown apps/navidrome/, apps/romm/
21:13 FINDING R-623 — the unattended caller turned every SUCCESS into a timeout Read before use, not after: follow() read update_phase/updating off the API ENVELOPE, so both were always None, every followed update hit the 900 s timeout and was then marked never-press-again. Fixed in that file before B1 relied on it. The earlier night missed it because its only follow() pass was the one whose log was lost update-arc-gaps-2026-09-21/unattended-caller.py